Deployment
Aegis speaks plain HTTP on port 2019 and never terminates TLS. Something has to sit in front of it. Here are three setups that work, and the details that bite.
Ready-to-run files for all of this live in deploy/ in the repository.
Three rules first
1. public_url is the contract
Every OAuth redirect URI, the discovery documents and the WWW-Authenticate challenge are derived from public_url. It must be the URL the client uses, including https:// — not the container name, not the internal address, not the IP.
Get this wrong and the symptom is confusing: the login page appears, you sign in, and the client then fails to exchange the code, because it was told to talk to a host it cannot reach.
2. Never publish port 2019
In every example below, the aegis service has no ports: section. Only the proxy reaches it, over the internal Docker network. Publishing 2019 on the host means a plaintext OAuth endpoint on the open internet, and browsers will send passwords to it.
3. Aegis trusts nothing from the proxy
It ignores X-Forwarded-Proto, X-Forwarded-Host and friends, and builds every URL from public_url instead. That makes it immune to header spoofing and means it needs no configuration to sit behind a proxy — with one consequence you should read: the client IP problem.
Traefik with Let's Encrypt
The standard case on your own server. Traefik terminates TLS, gets certificates automatically, and routes to aegis over the internal network.
name: aegis
services:
traefik:
image: traefik:v3.3
restart: unless-stopped
command:
- --providers.docker=true
- --providers.docker.exposedbydefault=false
- --entrypoints.web.address=:80
- --entrypoints.web.http.redirections.entrypoint.to=websecure
- --entrypoints.web.http.redirections.entrypoint.scheme=https
- --entrypoints.websecure.address=:443
- --certificatesresolvers.le.acme.email=${ACME_EMAIL}
- --certificatesresolvers.le.acme.storage=/letsencrypt/acme.json
- --certificatesresolvers.le.acme.tlschallenge=true
ports:
- "80:80"
- "443:443"
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
- letsencrypt:/letsencrypt
networks: [edge]
aegis:
image: ghcr.io/andreaskasper/aegis:latest
restart: unless-stopped
environment:
AEGIS_PUBLIC_URL: https://${AEGIS_HOST}
volumes:
- ./config.yaml:/etc/aegis/config.yaml:ro
read_only: true
cap_drop: [ALL]
security_opt:
- no-new-privileges:true
networks: [edge]
labels:
traefik.enable: "true"
traefik.http.routers.aegis.rule: "Host(`${AEGIS_HOST}`)"
traefik.http.routers.aegis.entrypoints: websecure
traefik.http.routers.aegis.tls.certresolver: le
traefik.http.services.aegis.loadbalancer.server.port: "2019"
traefik.http.services.aegis.loadbalancer.healthcheck.path: /healthz
traefik.http.services.aegis.loadbalancer.healthcheck.interval: 30s
networks:
edge:
volumes:
letsencrypt:
AEGIS_HOST=aegis.example.com ACME_EMAIL=you@example.com
Then:
cp ../../config.example.yaml config.yaml # and edit it docker compose up -d curl https://aegis.example.com/healthz
--certificatesresolvers.le.acme.caserver=https://acme-staging-v02.api.letsencrypt.org/directory
while you are getting DNS and labels right, then remove it and delete the letsencrypt volume once.
The Docker socket is mounted read-only. Traefik only needs to watch container labels; write access to that socket is equivalent to root on the host.
Cloudflare Tunnel
Nothing listens on the host at all. cloudflared opens an outbound connection to Cloudflare and requests arrive through it. No open port, no firewall exception, no certificate on the server. For a container that holds credentials, this is the smallest attack surface of the three.
name: aegis
services:
cloudflared:
image: cloudflare/cloudflared:latest
restart: unless-stopped
command: tunnel --no-autoupdate run
environment:
TUNNEL_TOKEN: ${TUNNEL_TOKEN}
networks: [internal]
depends_on: [aegis]
read_only: true
cap_drop: [ALL]
aegis:
image: ghcr.io/andreaskasper/aegis:latest
restart: unless-stopped
environment:
AEGIS_PUBLIC_URL: https://${AEGIS_HOST}
volumes:
- ./config.yaml:/etc/aegis/config.yaml:ro
read_only: true
cap_drop: [ALL]
security_opt:
- no-new-privileges:true
networks: [internal]
networks:
internal:
In the Zero Trust dashboard, under Networks → Tunnels → your tunnel → Public hostname:
| Field | Value |
|---|---|
| Subdomain / domain | aegis / example.com |
| Service type | HTTP |
| URL | aegis:2019 |
aegis is the compose service name; cloudflared resolves it over the internal Docker network. HTTP is correct here — that hop never leaves Docker, and TLS is terminated by Cloudflare at the edge.
.env with the file mode tightened, or hand it to the container as a Docker secret, and keep .env out of git.
Do not set internal: true on the network. That would also cut off aegis's own outbound requests to the APIs you configured, which is the entire point of the service.
Cloudflare's proxy in front of Traefik
The classic combination: the orange cloud at Cloudflare, your own server with Traefik and Let's Encrypt behind it. The compose file is the one from the Traefik section — what changes is the surrounding configuration.
| Setting | Value | Why |
|---|---|---|
| SSL/TLS mode | Full (strict) | Anything less lets Cloudflare talk to your origin over plaintext or an unverified certificate. With Let's Encrypt on Traefik, strict works out of the box. |
| DNS record | Proxied (orange) | If it is grey, you are in the plain Traefik setup and none of this applies. |
| ACME challenge | TLS-ALPN will fail | Cloudflare's proxy terminates TLS, so the TLS-ALPN challenge never reaches Traefik. Use the DNS-01 challenge, or issue the certificate once with the record set to grey and switch it back. |
| Caching | Leave default | Aegis sends Cache-Control: no-store on everything that matters. The discovery documents are deliberately cacheable for five minutes. |
| Cloudflare Access | See below | It will break MCP clients if applied naively. |
If you would rather not fight the ACME challenge, the honest alternative is a Cloudflare Origin Certificate: issue one in the dashboard, mount the certificate and key into Traefik as a static TLS store, and drop the certresolver entirely. It is valid only between Cloudflare and your origin, which is exactly what Full (strict) checks.
The client IP problem
This one is worth understanding before you rely on it.
Aegis rate-limits failed logins and client registrations per source IP, and it reads that IP from the TCP connection rather than from a forwarded header — deliberately, because a header an attacker controls is not an identity.
Behind a proxy, every connection comes from the proxy. So all of your users, and all of an attacker's attempts, land in the same bucket:
| Consequence | What it means for you |
|---|---|
| The limit becomes global | The default login_rate_limit: 10/m applies to the whole instance, not per person. |
| Lockout is possible | Someone hammering the login page can exhaust the bucket and stop you signing in, until it refills a few seconds later. |
| The audit log shows the proxy | login_failed lines carry the proxy's address, not the real client's. |
Practical mitigations, in the order I would apply them:
- Raise the limit to something that suits a shared bucket —
login_rate_limit: "60/m"— and rely on long passwords rather than on throttling. - Rate-limit at the proxy instead, where the real client IP is known. Traefik's
RateLimitmiddleware on the/authorizepath, or a Cloudflare WAF rate-limiting rule, does the job properly. - Put the login behind Cloudflare Access and exclude the machine-to-machine paths — see the next section.
X-Forwarded-For from a configured list of proxy addresses is on the roadmap; until then, the proxy is the right place to do per-IP limiting.
Cloudflare Access, and why it breaks things
Putting Cloudflare Access in front of the whole hostname is tempting and will stop your agent from working. Access expects a browser session; the MCP client is not a browser.
The flow splits into two kinds of request:
| Path | Who calls it | Access? |
|---|---|---|
/authorize | A human in a browser | Fine to protect |
/.well-known/* | The MCP client, unauthenticated | Must bypass |
/register | The MCP client | Must bypass |
/token | The MCP client | Must bypass |
/mcp | The MCP client, with a bearer token | Must bypass |
/healthz | Your monitoring | Must bypass |
So either protect nothing, or create an Access application scoped to /authorize only and leave the rest open. The machine paths are not unprotected in any meaningful sense: /mcp requires a bearer token, and /token requires a valid authorization code plus the PKCE verifier.
Adding Access on /authorize is genuinely useful, though — it puts your identity provider in front of the password form, which is the one endpoint where a human types a secret.
Timeouts and the SSE stream
Aegis offers GET /mcp as a Server-Sent Events stream. In v1 it carries nothing but a keep-alive comment every 30 seconds, so a proxy that closes it early is an annoyance rather than a failure — the client reconnects and everything still works over POST.
| Proxy | What to know |
|---|---|
| Traefik | Streams responses without buffering by default. Nothing to configure. |
| Cloudflare proxy | Passes SSE through, but idle connections are dropped after roughly 100 seconds. The 30-second keep-alive is below that, so the stream survives. |
| Cloudflare Tunnel | Same behaviour; no additional configuration. |
| nginx | Needs proxy_buffering off; for the SSE location, otherwise the stream is buffered and appears dead. |
The other timeout worth checking is the one on your upstreams. A slow API plus Cloudflare's 100-second edge limit means an http_request that takes longer will be cut off at the edge before aegis's own request_timeout fires. Keeping request_timeout at or below 30 seconds keeps aegis the one deciding.
Getting secrets into the container
Three ways, worst to best:
| Method | Config | Verdict |
|---|---|---|
Literal in config.yaml |
token: "sk-live-..." |
Works, but the file can never be committed and shows up in every backup. |
| Environment variable | token: "env:LEXWARE_TOKEN" |
Fine. Readable via docker inspect and in the process environment. |
| Docker secret | token: "file:/run/secrets/lexware" |
Best. Not in the environment, not in docker inspect, mounted as a tmpfs file. |
services:
aegis:
secrets:
- lexware
# ...
secrets:
lexware:
file: ./secrets/lexware_token
Then reference it as file:/run/secrets/lexware in config.yaml. Note that read_only: true does not interfere: Docker mounts secrets on their own tmpfs.
Validate before you start, so a typo does not become a restart loop:
docker compose run --rm aegis validate /etc/aegis/config.yaml
Updates and reloads
Two different operations, and only one of them costs you your sessions.
Changing the config
Edit config.yaml and aegis picks it up within two seconds; SIGHUP forces it immediately. A file that does not parse is rejected and the previous configuration keeps running. Sessions survive.
docker compose kill -s HUP aegis docker compose logs --tail 5 aegis
Changing the image
A restart wipes all in-memory state: every client re-registers and every user signs in again. That is by design, but it means you want to pin a version rather than have :latest move under you at an inconvenient moment.
# Pin to a digest and know exactly what is running docker compose pull docker compose up -d docker compose exec aegis /aegis version
Troubleshooting
| Symptom | Likely cause |
|---|---|
| Client says the server is not an OAuth server | The proxy is not routing /.well-known/*, or Access is intercepting it. Check with curl https://host/.well-known/oauth-authorization-server. |
| Login page loads, then the client fails | public_url does not match the hostname the client used. Every redirect is built from it. |
Redirect goes to http:// |
AEGIS_PUBLIC_URL is missing the https:// scheme. |
| 404 from the proxy | Traefik label loadbalancer.server.port is not 2019, or the container is not on the same network as Traefik. |
| 502 from the proxy | aegis exited. docker compose logs aegis — a config_error line at startup is the usual reason. |
403 origin not allowed |
A browser sent an Origin that is not public_url. Normal for a cross-origin experiment, never for a real client. |
| Everyone gets rate-limited at once | The client IP problem. |
short_secret warning at startup |
A secret under 8 characters cannot be redacted from responses. Replace it with a longer credential. |
When something is unclear, the audit log is the fastest way in — one JSON line per request, with the target, the status and the names of the secrets that were attached:
docker compose logs -f aegis | grep '"event":"request"'
aegis