Operations
Health checks, metrics, logging, indexes and a production checklist.
This page covers running ErtisAuth in production: health checks, monitoring, scaling, data maintenance and a security checklist.
Health checks#
GET /healthcheck#
Checks the database connection and whether the installation is set up. Anonymous.
| Situation | Status | Body |
|---|---|---|
| Database reachable, installation set up | 200 | { "status": "Healthy" } |
| Database reachable, not set up yet | 200 | { "status": "Unhealthy", "message": "ErtisAuth has not been set up yet" } |
| Database unreachable | 500 | { "status": "Unhealthy", "message": "Health check failed" } |
The details of a database failure are only written to the log.
200 even though it is Unhealthy. If your orchestrator only looks at the status code, it will treat a not-yet-set-up instance as healthy, which is what you want while you run the setup.GET /ping#
Answers "Pong" without touching the database. Use it as a liveness probe, and /healthcheck as a readiness probe.
livenessProbe:
httpGet: { path: /ping, port: 8080 }
readinessProbe:
httpGet: { path: /healthcheck, port: 8080 }Metrics#
Prometheus metrics are exposed at /metrics:
- incoming HTTP requests: counts, durations and in-progress requests per endpoint and status code,
- outgoing HTTP calls (to providers, webhook receivers, mail APIs),
- the .NET runtime: memory, garbage collection, thread pool.
/metrics is not authenticated. Don't expose it publicly: scrape it from inside your network, or block it at your ingress.
Logging and tracing#
ErtisAuth logs through the standard ASP.NET Core logging (console by default). Unexpected errors are logged with their stack trace, and the client only receives 500 UnhandledExceptionError.
When ApplicationInsights:ConnectionString is set, traces, metrics and logs are exported to Azure Monitor through OpenTelemetry (see Configuration).
API reference#
In the Development environment ErtisAuth serves its OpenAPI document at /openapi/v1.json and an interactive Scalar reference at /docs. Both are disabled in other environments.
Scaling#
ErtisAuth is stateless apart from its database, so you can run several instances behind a load balancer.
Caching#
Each instance keeps an in-memory cache of the resources it reads most often:
| Resource | Cached for up to |
|---|---|
| Memberships | 1 hour |
| User types | 1 hour |
| Providers | 1 hour |
| Roles | 5 minutes |
| Applications | 5 minutes |
| Revoked tokens (positive lookups only) | 24 hours |
A change is visible immediately on the instance that made it. Other instances see it when their cached copy expires. In practice:
- a role change takes up to 5 minutes to apply everywhere;
- a membership change (token lifetimes, mail providers, OTP settings…) takes up to 1 hour;
- revocations are seen everywhere immediately, since only revoked tokens are cached.
Restart the instances if a change must apply at once.
Background work#
Webhooks and mail hooks are processed from in-memory queues on the instance where the event occurred:
| Queue | Capacity | Parallel calls |
|---|---|---|
| Webhooks | 10,000 | 8 |
| Mails | 10,000 | 4 |
When a queue is full, new items are dropped and logged. On shutdown the queues are drained for up to 30 seconds; items still queued after that are lost. Give your pods a termination grace period of at least 30 seconds.
Database#
Indexes#
ErtisAuth creates the indexes it needs at startup, including:
- unique indexes for the unique fields of user types, kept in sync whenever a user type changes (and checked again at startup);
- text indexes for search;
- TTL indexes that delete expired data automatically: active tokens, revoked tokens, token codes and one-time passwords are removed a few minutes after they expire.
Retention#
The events collection is not cleaned up automatically and grows with every sign-in and change. Decide how long you need the audit log and remove older events regularly, for example with a TTL index of your own on event_time:
db.events.createIndex({ event_time: 1 }, { expireAfterSeconds: 60 * 60 * 24 * 180 })Backups#
Everything ErtisAuth knows is in its MongoDB database. Back it up like any production database, and protect the backups: they contain password hashes, membership secret keys, provider keys and mail provider credentials.
Security checklist#
- Serve ErtisAuth over HTTPS only. Basic tokens and passwords travel in plain text inside the TLS connection.
- Use a long random
secret_keyper membership (openssl rand -base64 48), and keep it secret. - Use
ARGON2IDas the hash algorithm for new memberships. - Keep
Database:ConnectionStringand other secrets out of source control. - Grant the sensitive read permissions only to operators:
memberships.readdiscloses secret keys and mail credentials,tokens.readdiscloses live tokens,events.readdiscloses user data.
- Give each application its own role with only what it needs, and rotate application secrets regularly.
- Rate-limit the public endpoints at your gateway:
/generate-token,/verify-otp,/oauth/{slug}/login,/memberships/{m}/codesand/memberships/{m}/codes/token. ErtisAuth doesn't limit request rates itself. - Restrict CORS at your gateway if your clients are only on known origins (ErtisAuth allows any origin).
- Don't log query strings of
/users/check-password(it carries a password). - Block
/metricsfrom the internet. - Set
trust_emailon providers only when you have checked what it means (see External Identity Providers). - Plan the retention of the
eventscollection.
Found a mistake in the docs? Open an issue