← Back to AI learnings
Deep Dive Aug 31, 2026

A task that reports success and writes nothing

AI Automation Claude

I run a set of scheduled Claude tasks that capture my sessions, audit my notes, and harvest writing ideas. For eight weeks they failed in three distinct ways, and I wrote down what each one taught me, because the lessons generalize past my setup.

The headline: a scheduled task that reports success and writes nothing is more dangerous than one that visibly crashes, because the failure stays invisible until something counts the missing files.

Failure mode one: silent success.

One weekly review task fired six times across five weeks and produced zero files. It reported success every single time. It took five weeks to notice, because "it ran" and "it worked" look identical from the outside.

The rule that falls out of this is short enough to tape to a monitor. A task whose output is a file is only healthy if the file exists. Check for the artifact, never the exit status.

Failure mode two: it dies mid-run and says nothing.

A Friday content harvest produced nothing for 12 consecutive weeks. For five of those weeks my diagnosis was timing, because it kept firing at odd hours. So I fixed the timing. Then it fired on time three Fridays in a row and still produced nothing.

When a fix does not change the outcome, the fix addressed a real problem that was not the problem. Stop fixing and go read the actual error. I spent five weeks not reading an error message.

Failure mode three: the whole batch fires late.

Three tasks on three different schedules all fired inside the same 0.4 seconds. That is not a scheduling bug. That is a wake-up batch, which means the question is about the machine, not the tasks. My laptop was asleep at the scheduled hour and everything caught up at once when it woke.

The ordering trap.

My pipeline has stages that depend on order: capture the week, then review it, then audit it. I spaced them by wall-clock time, half an hour apart. That is not ordering. That is a bet that nothing delays the batch. When the batch slipped, two reports that grade my capture system both ran before it and graded it on stale evidence. They were confidently wrong, in writing.

The observer was in the batch.

This is the one that stuck with me. The system watching for failures failed in the same way, at the same time, for the same reason. A quiet week and a broken robot look identical from the inside, and the thing meant to tell them apart was broken too.

The uncomfortable conclusion.

An automated maintenance system will endlessly repair the symptoms it has permission to touch while the one-line root-cause fix it cannot reach sits untouched. Mine recommended the same scheduler change seven weeks running and cheerfully fixed dozens of downstream symptoms in the meantime. Recommending the same fix an eighth time is not a strategy. Either widen the scope so the system can fix itself, or accept that symptom repair is the actual product you built.

And one last thing.

Two dead tasks came back to life with nothing edited. No change, no deploy, no explanation. Treat an unexplained recovery with exactly the same suspicion as an unexplained failure. One good week is not a fix.

If you're running scheduled AI tasks, the checklist is short. Alert on the missing artifact, not the exit code. Make dependent stages actually depend on each other instead of on the clock. Put the monitor somewhere the thing it monitors can't take down with it. And read the error before you fix the timing.