Scheduled Task Reliability
ActiveI run a set of scheduled Claude tasks that capture sessions, audit my notes, and harvest writing ideas. For eight weeks I watched them fail in three distinct ways and wrote down what each one taught me.
Silent success is the worst failure mode. One weekly review task fired six times over five weeks and wrote zero files. It reported success every single time. Detection took five weeks, because "it ran" and "it worked" look identical from the outside. The rule that falls out: a task whose output is a file is only healthy if the file exists. Check for the artifact, never the exit status.
When a fix doesn't change the outcome, you fixed a real problem that wasn't the problem. A Friday harvest went dark for 12 consecutive weeks. For five of those the diagnosis was timing. I fixed the timing. Three on-time Fridays later the output was still zero. Stop fixing and go read the actual error.
When several tasks with different schedules fire in the same second, that is not a scheduling bug. Three tasks on three different crons fired inside 0.4 seconds. That is a wake-up batch, which means the question is about the machine, not the tasks. My laptop was asleep at the scheduled hour.
The ordering trap. If pipeline stages depend on order, spacing them by wall-clock time is not ordering. It is a bet that nothing delays the batch. When the batch slipped, two reports that grade my capture system both ran before it and graded it on stale evidence.
The observer was in the batch. The system watching for failures failed in the same way, at the same time, for the same reason. A quiet week and a broken robot look identical from the inside.
The uncomfortable one. An automated maintenance system will endlessly repair symptoms it has permission to touch while the one-line root-cause fix it cannot reach sits untouched. Recommending the same fix an eighth time is not a strategy. Either widen the scope so the system can fix itself, or accept that symptom repair is the actual product.
And one more: treat an unexplained recovery with the same suspicion as an unexplained failure. Two dead tasks came back to life with nothing edited. One working week is not a fix.
Still running, still not fully fixed. That is the honest status.