A URL
Warden loads it on a clock. Two failures in a row open a problem; one blip does not.
When your app goes down, Warden works out why, brings it back the way you allowed, checks it is really back, and tells you. You hear from it only when the decision is yours.
Warden loads it on a clock. Two failures in a row open a problem; one blip does not.
Render, Railway, Vercel and Coolify each give you a URL that redeploys the app when POSTed. Warden calls it when your rules allow, and only then.
A Slack or Discord webhook. You hear when something breaks and when Warden stops to ask. Nothing else.
At 3am, from a phone, a person does the same four things in the same order. An agent with a shell on your production box is a worse problem than the outage, so Warden has no shell: it picks one action by name from a fixed list of 20, the arguments are validated, and it is spawned without a shell.
The process table, the logs, what landed recently, the diff of the one commit that looks relevant. Ten looks, and an investigation that keeps reading is avoiding a conclusion.
In plain words, quoting what it actually read, with a number for how sure it is. Under fifty per cent it does not get to act at all.
Not the agent. A pure function reads the rules you wrote for that service and answers with the one that decided: do it, refuse it, or stop the run and ask you.
Warden re-runs the exact check that failed. A clean reading closes the problem and its id is stored as the proof. Warden never gets to say it fixed something.
Two boundaries, both in code, neither of which the model can move. Everything else about this product is downstream of them.
One set per service, written by a human, action by action. May — it does it and tells you afterwards. Ask — it works out exactly what it would do, then stops the run and waits, however long that takes. Never — refused, with the rule that refused it named. Plus a cap per problem and a cooldown per service, because a granted permission is not an unbounded one.
decide(op, policy, ctx) → { verdict, rule, reason }
// pure. no clock, no model, no network.A problem is closed by one thing and it is not the agent’s opinion: the same check that opened it, run again, in code. The reading’s id is stored on the problem. If it comes back failing, the problem says “Warden acted, but the check still fails” — and you are woken.
resolveIncident(id, reading, resolution) // verifiedByReadingId — the proof, not a claim.
Four of the 20 are declared forbidden rather than left out — migrating a database, deleting data, rotating a secret, destroying infrastructure — so the product can show you the line. No rule can turn them on. See all 20 against a real set of rules →
A working operator has a boring board, which is a genuine presentation problem, and the honest answer to it is a real outage rather than a video. Press it and Warden’s own checks notice, a problem opens, and you land on it with the run already streaming.
“vigil is down because the process is stopped … the error log is empty and stdout shows only clean startup banners. So I cannot name the trigger of the stop from logs alone — it was silent. … A restart is the right next act; if it dies again immediately with no logged error, look at commits 6fad866 and 503accd — they bracket when this started getting unhealthy.”
a real diagnosis, from the audit table · the rules allowed a restart · the check came back 200 in 207ms · down 60s
No account, no email, nothing to install. Give Warden something you have running and it starts checking. Tell it the machine as well and it can read the logs, the process table and the last few commits when that URL stops answering — and, if you let it, put the thing back up and prove it.