IT Buddy
← Back to blog
AI Implementation 7 min read

The night watch: how we let Claude work when nobody is watching

Uros Vujic 30. september 2026

The one who isn't watching

This is the third post in a series. In the first we wrote about the loop: that most of the progress in AI now comes from the way of working around the model. The second was about the stopping condition, meaning how the loop knows it is finished.

This post is about the next step. What happens when the loop runs and nobody is watching?

We have worked this way on one of our own projects over the past few weeks. We call it Nattravn here, Norwegian for nightjar, because the project itself is not the point. The plans Claude works from we call night watch plans. Not because everything happens at night, but because they are written for work that nobody supervises. Whether it is night or afternoon matters less. What matters is that the plan has to cope with you not being there.

For me, it has considerably increased how much I get done. But that did not happen by itself. It came from the rules in the plan.

What a night watch plan is

A night watch plan splits the work into small tasks, and each task passes through three roles:

The one who does the work. One agent gets one task with a precise description. Where the task has test steps, it writes the test first, sees it fail for the right reason, builds the solution, and sees the test pass. Then it commits and writes a report.

The one who checks. A separate agent reads the change and compares it with the task. It has one clear mandate: the report from the one who did the work is to be treated as unverified claims.

The one who steers. Keeps the overview, sends the next task, and keeps a log of everything that happens.

Every task has to end with one of four answers:

Answer Means
Done The task is finished and tested
Done, with concerns Finished, but something deviates from the plan and needs a look
Blocked Cannot go further, and says why
Needs more information Something is unclear, and it does not guess

The last one matters more than it looks. An agent without supervision that guesses will guess for hours.

Five rules we have learned

1. Stop rather than guess

The plan says it plainly: bad work is worse than no work.

That was put to the test. Midway through Nattravn, an agent took eight screenshots that were meant to be used further on. All eight showed a red quality score of 0 per cent, because the demo data was incomplete. Fixing it required changes outside the project the agent was allowed to touch.

The one steering stopped and asked, instead of delivering something that looked finished. That is exactly what a night watch should do.

2. The one who does the work should not approve it

This is the rule with the most outside support. The Claude Code documentation recommends a separate verification agent, so that the agent doing the work is not the one grading it.

A research paper from July shows why. Agents that assessed their own progress reported improvement in every single round. In 56 per cent of the rounds, the measured change was zero or negative. When the agent itself was allowed to let changes through, it eventually accepted everything, and the result ended up 19 per cent below the best it had already reached.

In Nattravn, the checker found twice that the report from the one who did the work was wrong about the number of tests and lines. The code was right. The report was not. It is a small example, but it is exactly what the paper describes.

And the check is not infallible either. Once, an image with legible text slipped through where there should have been none. It was caught at the next step. That is why the plan also has gates where a person has to approve before the work moves on.

3. Write down every decision, and what it costs if it is wrong

The one steering sometimes has to make decisions on its own. In Nattravn there were eleven. Each one is logged with a single sentence at the end: what it costs if it is wrong.

"A one-line change." "None beyond a branch rename." "A soft crop edge visible in one 4 s scene."

That sentence is the whole point. The agent is allowed to decide alone when being wrong is cheap. When being wrong is expensive, the decision is yours.

4. Set limits before you leave

A loop without limits does not stop by itself. The night watch plans have three kinds:

A limit on attempts. Each task gets at most five fix rounds. In Nattravn, seven were used in total, spread across all the tasks.

A limit on money. Where the work costs money, the price is checked before every job, and there is a cap both per task and for the whole project.

Things it may never do. It never commits on the main branch. The code of the product itself is read-only. And Claude never types my password. I create the login myself, and until it exists, the task waits.

Anthropic describes the same thing from its own experiments: agents without such limits tried to do everything at once and declared themselves done before they were. Their solution is one task at a time, a progress log, and a commit after every task.

5. Change course with a new job, not a message mid-run

We learned this one the hard way. In the middle of a task, a message was sent asking it to add something and stop. The agent treated the message as a possible attempt at manipulation, ignored it, and finished the original task.

That is actually a good property. An agent without supervision should not obey whatever turns up along the way. The lesson in the log is simple: if you want to change course, send a new task.

The small things that make a big difference

Two more things are worth including, even though they are not rules.

Check before you start. Before the first task runs, the plan compares what each task produces with what the next ones need. In Nattravn, that check found a conflict between a colour rule and a task before anything had been built.

Park the small stuff. The checks found 18 minor issues along the way. None of them was allowed to stop the work. They were logged and dealt with later. Without that, a night watch gets sidetracked by details.

And git is the safety net. Once, a session was interrupted mid-work and left the repository in a broken state. Because everything had been committed task by task, it was restored without any work being lost.

What it actually came to

From 25 to 28 September, Nattravn completed 26 tasks and 53 commits, with 20 separate review rounds along the way.

My role in the logs is not writing code. It is approving at the gates, choosing between alternatives, and making the decisions that are too expensive for the agent to make alone. That is a different job from sitting next to it and watching. And it is that job that lets the rest run without me.

How to get started

You do not need a big project to try this. You need five things in place before you walk away:

A task small enough that it can be checked on its own.

A check it can run itself, such as a test, a build or a screenshot to compare against.

A checker that is not the same as the one who did the work.

Limits on attempts, money, and what it may never do.

Permission to stop. Write it into the plan: better blocked than wrong.

The rest is discipline. And it is the same discipline as always: structure first, technology after.

UV

Uros Vujic

Daglig leder, IT Buddy AS

Uros hjelper norske virksomheter med å innføre AI på en kontrollert og bærekraftig måte. Bakgrunn fra IT-infrastruktur i bank og finans, med spesialisering i AI governance, RBAC og GDPR-compliant implementering.

Ready for the next step?

Take our AI Ready assessment and find out where your business stands.

Take AI Ready Assessment