Deployment

Aegis speaks plain HTTP on port 2019 and never terminates TLS. Something has to sit in front of it. Here are three setups that work, and the details that bite.

Ready-to-run files for all of this live in deploy/ in the repository.

Three rules first

1. public_url is the contract

Every OAuth redirect URI, the discovery documents and the WWW-Authenticate challenge are derived from public_url. It must be the URL the client uses, including https:// — not the container name, not the internal address, not the IP.

Get this wrong and the symptom is confusing: the login page appears, you sign in, and the client then fails to exchange the code, because it was told to talk to a host it cannot reach.

2. Never publish port 2019

In every example below, the aegis service has no ports: section. Only the proxy reaches it, over the internal Docker network. Publishing 2019 on the host means a plaintext OAuth endpoint on the open internet, and browsers will send passwords to it.

3. Aegis trusts nothing from the proxy

It ignores X-Forwarded-Proto, X-Forwarded-Host and friends, and builds every URL from public_url instead. That makes it immune to header spoofing and means it needs no configuration to sit behind a proxy — with one consequence you should read: the client IP problem.

Traefik with Let's Encrypt

The standard case on your own server. Traefik terminates TLS, gets certificates automatically, and routes to aegis over the internal network.

deploy/traefik/docker-compose.yml
name: aegis

services:
  traefik:
    image: traefik:v3.3
    restart: unless-stopped
    command:
      - --providers.docker=true
      - --providers.docker.exposedbydefault=false
      - --entrypoints.web.address=:80
      - --entrypoints.web.http.redirections.entrypoint.to=websecure
      - --entrypoints.web.http.redirections.entrypoint.scheme=https
      - --entrypoints.websecure.address=:443
      - --certificatesresolvers.le.acme.email=${ACME_EMAIL}
      - --certificatesresolvers.le.acme.storage=/letsencrypt/acme.json
      - --certificatesresolvers.le.acme.tlschallenge=true
    ports:
      - "80:80"
      - "443:443"
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - letsencrypt:/letsencrypt
    networks: [edge]

  aegis:
    image: ghcr.io/andreaskasper/aegis:latest
    restart: unless-stopped
    environment:
      AEGIS_PUBLIC_URL: https://${AEGIS_HOST}
    volumes:
      - ./config.yaml:/etc/aegis/config.yaml:ro
    read_only: true
    cap_drop: [ALL]
    security_opt:
      - no-new-privileges:true
    networks: [edge]
    labels:
      traefik.enable: "true"
      traefik.http.routers.aegis.rule: "Host(`${AEGIS_HOST}`)"
      traefik.http.routers.aegis.entrypoints: websecure
      traefik.http.routers.aegis.tls.certresolver: le
      traefik.http.services.aegis.loadbalancer.server.port: "2019"
      traefik.http.services.aegis.loadbalancer.healthcheck.path: /healthz
      traefik.http.services.aegis.loadbalancer.healthcheck.interval: 30s

networks:
  edge:

volumes:
  letsencrypt:
.env
AEGIS_HOST=aegis.example.com
ACME_EMAIL=you@example.com

Then:

shell
cp ../../config.example.yaml config.yaml   # and edit it
docker compose up -d
curl https://aegis.example.com/healthz
Test against the staging CA first. Let's Encrypt rate-limits failed attempts hard. Add --certificatesresolvers.le.acme.caserver=https://acme-staging-v02.api.letsencrypt.org/directory while you are getting DNS and labels right, then remove it and delete the letsencrypt volume once.

The Docker socket is mounted read-only. Traefik only needs to watch container labels; write access to that socket is equivalent to root on the host.

Cloudflare Tunnel

Nothing listens on the host at all. cloudflared opens an outbound connection to Cloudflare and requests arrive through it. No open port, no firewall exception, no certificate on the server. For a container that holds credentials, this is the smallest attack surface of the three.

deploy/cloudflared/docker-compose.yml
name: aegis

services:
  cloudflared:
    image: cloudflare/cloudflared:latest
    restart: unless-stopped
    command: tunnel --no-autoupdate run
    environment:
      TUNNEL_TOKEN: ${TUNNEL_TOKEN}
    networks: [internal]
    depends_on: [aegis]
    read_only: true
    cap_drop: [ALL]

  aegis:
    image: ghcr.io/andreaskasper/aegis:latest
    restart: unless-stopped
    environment:
      AEGIS_PUBLIC_URL: https://${AEGIS_HOST}
    volumes:
      - ./config.yaml:/etc/aegis/config.yaml:ro
    read_only: true
    cap_drop: [ALL]
    security_opt:
      - no-new-privileges:true
    networks: [internal]

networks:
  internal:

In the Zero Trust dashboard, under Networks → Tunnels → your tunnel → Public hostname:

FieldValue
Subdomain / domainaegis / example.com
Service typeHTTP
URLaegis:2019

aegis is the compose service name; cloudflared resolves it over the internal Docker network. HTTP is correct here — that hop never leaves Docker, and TLS is terminated by Cloudflare at the edge.

The tunnel token is a credential. Anyone holding it can attach a connector to your tunnel. Keep it in .env with the file mode tightened, or hand it to the container as a Docker secret, and keep .env out of git.

Do not set internal: true on the network. That would also cut off aegis's own outbound requests to the APIs you configured, which is the entire point of the service.

Cloudflare's proxy in front of Traefik

The classic combination: the orange cloud at Cloudflare, your own server with Traefik and Let's Encrypt behind it. The compose file is the one from the Traefik section — what changes is the surrounding configuration.

SettingValueWhy
SSL/TLS mode Full (strict) Anything less lets Cloudflare talk to your origin over plaintext or an unverified certificate. With Let's Encrypt on Traefik, strict works out of the box.
DNS record Proxied (orange) If it is grey, you are in the plain Traefik setup and none of this applies.
ACME challenge TLS-ALPN will fail Cloudflare's proxy terminates TLS, so the TLS-ALPN challenge never reaches Traefik. Use the DNS-01 challenge, or issue the certificate once with the record set to grey and switch it back.
Caching Leave default Aegis sends Cache-Control: no-store on everything that matters. The discovery documents are deliberately cacheable for five minutes.
Cloudflare Access See below It will break MCP clients if applied naively.

If you would rather not fight the ACME challenge, the honest alternative is a Cloudflare Origin Certificate: issue one in the dashboard, mount the certificate and key into Traefik as a static TLS store, and drop the certresolver entirely. It is valid only between Cloudflare and your origin, which is exactly what Full (strict) checks.

The client IP problem

This one is worth understanding before you rely on it.

Aegis rate-limits failed logins and client registrations per source IP, and it reads that IP from the TCP connection rather than from a forwarded header — deliberately, because a header an attacker controls is not an identity.

Behind a proxy, every connection comes from the proxy. So all of your users, and all of an attacker's attempts, land in the same bucket:

ConsequenceWhat it means for you
The limit becomes globalThe default login_rate_limit: 10/m applies to the whole instance, not per person.
Lockout is possibleSomeone hammering the login page can exhaust the bucket and stop you signing in, until it refills a few seconds later.
The audit log shows the proxylogin_failed lines carry the proxy's address, not the real client's.

Practical mitigations, in the order I would apply them:

  • Raise the limit to something that suits a shared bucket — login_rate_limit: "60/m" — and rely on long passwords rather than on throttling.
  • Rate-limit at the proxy instead, where the real client IP is known. Traefik's RateLimit middleware on the /authorize path, or a Cloudflare WAF rate-limiting rule, does the job properly.
  • Put the login behind Cloudflare Access and exclude the machine-to-machine paths — see the next section.
This is a known limitation, not a mystery. Teaching aegis to trust X-Forwarded-For from a configured list of proxy addresses is on the roadmap; until then, the proxy is the right place to do per-IP limiting.

Cloudflare Access, and why it breaks things

Putting Cloudflare Access in front of the whole hostname is tempting and will stop your agent from working. Access expects a browser session; the MCP client is not a browser.

The flow splits into two kinds of request:

PathWho calls itAccess?
/authorizeA human in a browserFine to protect
/.well-known/*The MCP client, unauthenticatedMust bypass
/registerThe MCP clientMust bypass
/tokenThe MCP clientMust bypass
/mcpThe MCP client, with a bearer tokenMust bypass
/healthzYour monitoringMust bypass

So either protect nothing, or create an Access application scoped to /authorize only and leave the rest open. The machine paths are not unprotected in any meaningful sense: /mcp requires a bearer token, and /token requires a valid authorization code plus the PKCE verifier.

Adding Access on /authorize is genuinely useful, though — it puts your identity provider in front of the password form, which is the one endpoint where a human types a secret.

Timeouts and the SSE stream

Aegis offers GET /mcp as a Server-Sent Events stream. In v1 it carries nothing but a keep-alive comment every 30 seconds, so a proxy that closes it early is an annoyance rather than a failure — the client reconnects and everything still works over POST.

ProxyWhat to know
TraefikStreams responses without buffering by default. Nothing to configure.
Cloudflare proxyPasses SSE through, but idle connections are dropped after roughly 100 seconds. The 30-second keep-alive is below that, so the stream survives.
Cloudflare TunnelSame behaviour; no additional configuration.
nginxNeeds proxy_buffering off; for the SSE location, otherwise the stream is buffered and appears dead.

The other timeout worth checking is the one on your upstreams. A slow API plus Cloudflare's 100-second edge limit means an http_request that takes longer will be cut off at the edge before aegis's own request_timeout fires. Keeping request_timeout at or below 30 seconds keeps aegis the one deciding.

Getting secrets into the container

Three ways, worst to best:

MethodConfigVerdict
Literal in config.yaml token: "sk-live-..." Works, but the file can never be committed and shows up in every backup.
Environment variable token: "env:LEXWARE_TOKEN" Fine. Readable via docker inspect and in the process environment.
Docker secret token: "file:/run/secrets/lexware" Best. Not in the environment, not in docker inspect, mounted as a tmpfs file.
docker-compose.yml (secrets)
services:
  aegis:
    secrets:
      - lexware
    # ...

secrets:
  lexware:
    file: ./secrets/lexware_token

Then reference it as file:/run/secrets/lexware in config.yaml. Note that read_only: true does not interfere: Docker mounts secrets on their own tmpfs.

Validate before you start, so a typo does not become a restart loop:

shell
docker compose run --rm aegis validate /etc/aegis/config.yaml

Updates and reloads

Two different operations, and only one of them costs you your sessions.

Changing the config

Edit config.yaml and aegis picks it up within two seconds; SIGHUP forces it immediately. A file that does not parse is rejected and the previous configuration keeps running. Sessions survive.

shell
docker compose kill -s HUP aegis
docker compose logs --tail 5 aegis

Changing the image

A restart wipes all in-memory state: every client re-registers and every user signs in again. That is by design, but it means you want to pin a version rather than have :latest move under you at an inconvenient moment.

shell
# Pin to a digest and know exactly what is running
docker compose pull
docker compose up -d
docker compose exec aegis /aegis version

Troubleshooting

SymptomLikely cause
Client says the server is not an OAuth server The proxy is not routing /.well-known/*, or Access is intercepting it. Check with curl https://host/.well-known/oauth-authorization-server.
Login page loads, then the client fails public_url does not match the hostname the client used. Every redirect is built from it.
Redirect goes to http:// AEGIS_PUBLIC_URL is missing the https:// scheme.
404 from the proxy Traefik label loadbalancer.server.port is not 2019, or the container is not on the same network as Traefik.
502 from the proxy aegis exited. docker compose logs aegis — a config_error line at startup is the usual reason.
403 origin not allowed A browser sent an Origin that is not public_url. Normal for a cross-origin experiment, never for a real client.
Everyone gets rate-limited at once The client IP problem.
short_secret warning at startup A secret under 8 characters cannot be redacted from responses. Replace it with a longer credential.

When something is unclear, the audit log is the fastest way in — one JSON line per request, with the target, the status and the names of the secrets that were attached:

shell
docker compose logs -f aegis | grep '"event":"request"'