Watching 3 real services right now

Your software should not need you awake
to keep running.

When your app goes down, Warden works out why, brings it back the way you allowed, checks it is really back, and tells you. You hear from it only when the decision is yours.

3,600checks run
33problems
31closed without waking anyone
2mmedian time down
The fleet, right now live
upWarden's own consoleasks first
200 in 673ms
2 openVigilmay act
expected 200, got 502
upSAGEobserve only
200 in 233ms
  • checked SAGE · 56 days left, until 2026-12-06 (Let's Encrypt) in 12ms
  • checked SAGE · online, 16 restarts since deploy in 227ms
  • checked SAGE · 200 in 233ms
  • checked Vigil · 63 days left, until 2026-12-13 (Let's Encrypt) in 13ms
  • checked Vigil · pm2 says "stopped" in 184ms
  • checked Vigil · expected 200, got 502 in 710ms
last checked …3 services
All it needs

A URL, your deploy hook, and where to reach you.

A URL

Warden loads it on a clock. Two failures in a row open a problem; one blip does not.

A deploy hook

Render, Railway, Vercel and Coolify each give you a URL that redeploys the app when POSTed. Warden calls it when your rules allow, and only then.

Where to reach you

A Slack or Discord webhook. You hear when something breaks and when Warden stops to ask. Nothing else.

The twenty minutes after the alert

The work is mechanical. That is why it is worth automating — and why it is frightening.

At 3am, from a phone, a person does the same four things in the same order. An agent with a shell on your production box is a worse problem than the outage, so Warden has no shell: it picks one action by name from a fixed list of 20, the arguments are validated, and it is spawned without a shell.

  1. 01

    It looks

    The process table, the logs, what landed recently, the diff of the one commit that looks relevant. Ten looks, and an investigation that keeps reading is avoiding a conclusion.

  2. 02

    It commits to a cause

    In plain words, quoting what it actually read, with a number for how sure it is. Under fifty per cent it does not get to act at all.

  3. 03

    Your rules decide

    Not the agent. A pure function reads the rules you wrote for that service and answers with the one that decided: do it, refuse it, or stop the run and ask you.

  4. 04

    The check decides whether it worked

    Warden re-runs the exact check that failed. A clean reading closes the problem and its id is stored as the proof. Warden never gets to say it fixed something.

What it may do, and how it proves it

Give it permission to act. Not permission to do anything.

Two boundaries, both in code, neither of which the model can move. Everything else about this product is downstream of them.

Your rules decide whether an act happens

One set per service, written by a human, action by action. May — it does it and tells you afterwards. Ask — it works out exactly what it would do, then stops the run and waits, however long that takes. Never — refused, with the rule that refused it named. Plus a cap per problem and a cooldown per service, because a granted permission is not an unbounded one.

decide(op, policy, ctx) → { verdict, rule, reason }
// pure. no clock, no model, no network.

The check decides whether it worked

A problem is closed by one thing and it is not the agent’s opinion: the same check that opened it, run again, in code. The reading’s id is stored on the problem. If it comes back failing, the problem says “Warden acted, but the check still fails” — and you are woken.

resolveIncident(id, reading, resolution)
// verifiedByReadingId — the proof, not a claim.
check the certificateReads the TLS certificate a URL is served with and reports how many days are left on it.looks only
restart the processRestarts the process. Undoes itself: the same code comes back up.undoes itself
roll back to the last deployChecks out the previous commit and restarts. This changes which code is running.disruptive
delete dataWould delete data. Refused by name. There is no incident this is the answer to.never, for anyone
destroy infrastructureWould destroy infrastructure. Refused by name, and listed here so you can see it is.never, for anyone

Four of the 20 are declared forbidden rather than left out — migrating a database, deleting data, rotating a secret, destroying infrastructure — so the product can show you the line. No rule can turn them on. See all 20 against a real set of rules →

Not a recording

There is a button on the board that really stops a real service.

A working operator has a boring board, which is a genuine presentation problem, and the honest answer to it is a real outage rather than a video. Press it and Warden’s own checks notice, a problem opens, and you land on it with the run already streaming.

31problems closed without waking anyone
2mmedian time a service was down
3,600checks written down, including the boring ones
327tests, offline — no network, no model, no API key

“vigil is down because the process is stopped … the error log is empty and stdout shows only clean startup banners. So I cannot name the trigger of the stop from logs alone — it was silent. … A restart is the right next act; if it dies again immediately with no logged error, look at commits 6fad866 and 503accd — they bracket when this started getting unhealthy.”

a real diagnosis, from the audit table · the rules allowed a restart · the check came back 200 in 207ms · down 60s

A URL is the whole sign-up.

No account, no email, nothing to install. Give Warden something you have running and it starts checking. Tell it the machine as well and it can read the logs, the process table and the last few commits when that URL stops answering — and, if you let it, put the thing back up and prove it.

Warden — an autonomous operator for software that is already running