ErtisAuth

Operations

Health checks, metrics, logging, indexes and a production checklist.

This page covers running ErtisAuth in production: health checks, monitoring, scaling, data maintenance and a security checklist.

Health checks#

GET /healthcheck#

Checks the database connection and whether the installation is set up. Anonymous.

SituationStatusBody
Database reachable, installation set up200{ "status": "Healthy" }
Database reachable, not set up yet200{ "status": "Unhealthy", "message": "ErtisAuth has not been set up yet" }
Database unreachable500{ "status": "Unhealthy", "message": "Health check failed" }

The details of a database failure are only written to the log.

Note: a fresh installation answers 200 even though it is Unhealthy. If your orchestrator only looks at the status code, it will treat a not-yet-set-up instance as healthy, which is what you want while you run the setup.

GET /ping#

Answers "Pong" without touching the database. Use it as a liveness probe, and /healthcheck as a readiness probe.

yaml
livenessProbe:
  httpGet: { path: /ping, port: 8080 }
readinessProbe:
  httpGet: { path: /healthcheck, port: 8080 }

Metrics#

Prometheus metrics are exposed at /metrics:

  • incoming HTTP requests: counts, durations and in-progress requests per endpoint and status code,
  • outgoing HTTP calls (to providers, webhook receivers, mail APIs),
  • the .NET runtime: memory, garbage collection, thread pool.

/metrics is not authenticated. Don't expose it publicly: scrape it from inside your network, or block it at your ingress.

Logging and tracing#

ErtisAuth logs through the standard ASP.NET Core logging (console by default). Unexpected errors are logged with their stack trace, and the client only receives 500 UnhandledExceptionError.

When ApplicationInsights:ConnectionString is set, traces, metrics and logs are exported to Azure Monitor through OpenTelemetry (see Configuration).

API reference#

In the Development environment ErtisAuth serves its OpenAPI document at /openapi/v1.json and an interactive Scalar reference at /docs. Both are disabled in other environments.

Scaling#

ErtisAuth is stateless apart from its database, so you can run several instances behind a load balancer.

Caching#

Each instance keeps an in-memory cache of the resources it reads most often:

ResourceCached for up to
Memberships1 hour
User types1 hour
Providers1 hour
Roles5 minutes
Applications5 minutes
Revoked tokens (positive lookups only)24 hours

A change is visible immediately on the instance that made it. Other instances see it when their cached copy expires. In practice:

  • a role change takes up to 5 minutes to apply everywhere;
  • a membership change (token lifetimes, mail providers, OTP settings…) takes up to 1 hour;
  • revocations are seen everywhere immediately, since only revoked tokens are cached.

Restart the instances if a change must apply at once.

Background work#

Webhooks and mail hooks are processed from in-memory queues on the instance where the event occurred:

QueueCapacityParallel calls
Webhooks10,0008
Mails10,0004

When a queue is full, new items are dropped and logged. On shutdown the queues are drained for up to 30 seconds; items still queued after that are lost. Give your pods a termination grace period of at least 30 seconds.

Database#

Indexes#

ErtisAuth creates the indexes it needs at startup, including:

  • unique indexes for the unique fields of user types, kept in sync whenever a user type changes (and checked again at startup);
  • text indexes for search;
  • TTL indexes that delete expired data automatically: active tokens, revoked tokens, token codes and one-time passwords are removed a few minutes after they expire.

Retention#

The events collection is not cleaned up automatically and grows with every sign-in and change. Decide how long you need the audit log and remove older events regularly, for example with a TTL index of your own on event_time:

javascript
db.events.createIndex({ event_time: 1 }, { expireAfterSeconds: 60 * 60 * 24 * 180 })

Backups#

Everything ErtisAuth knows is in its MongoDB database. Back it up like any production database, and protect the backups: they contain password hashes, membership secret keys, provider keys and mail provider credentials.

Security checklist#

  • Serve ErtisAuth over HTTPS only. Basic tokens and passwords travel in plain text inside the TLS connection.
  • Use a long random secret_key per membership (openssl rand -base64 48), and keep it secret.
  • Use ARGON2ID as the hash algorithm for new memberships.
  • Keep Database:ConnectionString and other secrets out of source control.
  • Grant the sensitive read permissions only to operators:
    • memberships.read discloses secret keys and mail credentials,
    • tokens.read discloses live tokens,
    • events.read discloses user data.
  • Give each application its own role with only what it needs, and rotate application secrets regularly.
  • Rate-limit the public endpoints at your gateway: /generate-token, /verify-otp, /oauth/{slug}/login, /memberships/{m}/codes and /memberships/{m}/codes/token. ErtisAuth doesn't limit request rates itself.
  • Restrict CORS at your gateway if your clients are only on known origins (ErtisAuth allows any origin).
  • Don't log query strings of /users/check-password (it carries a password).
  • Block /metrics from the internet.
  • Set trust_email on providers only when you have checked what it means (see External Identity Providers).
  • Plan the retention of the events collection.

Found a mistake in the docs? Open an issue