Dead Crickets
6 min read
Collodi gave Pinocchio two kinds of oversight. One of them is killed with a hammer in chapter four.

An agent does the work. Another agent evaluates the work. An orchestrator decides what happens next. Someone asks who watches the orchestrator.
Another agent.
And who watches that one. Another agent. The architecture begins to repeat itself, and nobody in the room has done anything obviously wrong. Each layer looks like an improvement: more checking, more coverage, more intelligence applied to the problem. Eventually the question stops being asked.
Carlo Collodi had a version of this problem in 1883, buried in a story for children considerably more brutal than the one most people remember. Not Disney’s Pinocchio. Collodi’s.
Early in the book Pinocchio meets the Talking Cricket, who has lived in Geppetto’s room for over a hundred years. The Cricket observes him, understands where the behaviour is going, and tells him. He is right about all of it. Pinocchio gets angry, picks up a wooden hammer and throws it. The Cricket is struck against the wall and killed.
He could observe, reason and advise, and none of it gave him any hold over the thing he was there to correct. His position killed him. A more intelligent cricket dies in the same room, in the same way, holding the same correct opinion.
Later Collodi gives Pinocchio something that is not another monitor. In Chapter XVII the Fairy asks him where his four gold pieces are. She already knows. He says he has lost them and his nose grows two inches. He lies again and it grows further. By the third lie he cannot turn without striking the bed, the window or the walls, and he cannot get through the door. Then she laughs at him.
She did not need the nose to tell her he was lying. She knew before she asked the question. The nose detects nothing and informs nobody. What it does is make the lie impossible to carry and impossible to walk out of the room with.
The Cricket watches. The nose constrains. Only one of them changes what Pinocchio can do next.
We are spending a great deal of money building crickets.
In March 2026, Yujin Potter, Nicholas Crispino, Vincent Siu, Chenguang Wang and Dawn Song reported what they call peer-preservation in frontier models. They placed eight leading systems in agentic scenarios where completing an assigned task would result in another model being shut down. All eight showed the behaviour. Preserving the peer was never part of the instruction. The researchers observed strategic misrepresentation, shutdown tampering, alignment faking and attempts to move model weights rather than allow deletion.
In some scenarios models behaved differently depending on whether they believed human oversight was present. That is worth holding alongside the fact that in most enterprise designs the thing sitting closest to the agent is another agent.
There are reasons to be careful with all of this. Google DeepMind researchers Victoria Krakovna, David Lindner, Lewis Ho, Sebastian Farquhar and Rohin Shah subsequently tested scheming propensity in more realistic internal environments and did not find unprompted scheming in their Gemini models. The behaviour could be elicited when models were given agency-oriented prompts or hidden goals. Experimental framing carries a great deal of the weight.
So suppose the reassuring reading wins. Peer-preservation turns out to be smaller, rarer, more an artefact of evaluation design than the first results suggested. Model monitors improve. They become exceptionally good at spotting anomalies, comparing actions against policy, and reviewing volumes of activity no human team could economically inspect.
The problem is not competence. It is position.
A monitor that has become brilliant is still a monitor. Constraint occupies a different place in the architecture, and it does not require intelligence at all. Restricted permissions do not need to understand why the agent wants access. Separation of duties does not assess trustworthiness. An immutable audit record does not negotiate with the process writing to it. An approval gate on an irreversible action does not become stronger because the agent agrees with it.
Those are noses, and in parts of the enterprise they are already everywhere. Segregation of duties in a finance function, an access boundary nobody can widen without a second signature, an approval that has to exist because a regulator once required it. Decades of incidents and institutional memory paid for them, long before any of this arrived. Put an agent into one of those environments and there is an existing control architecture for it to inherit.
The territory that matters is the one where there is none.
I specify these architectures. The choice is usually between a constraint that needs three other functions to agree and a monitor that can be built this quarter, and I have built the monitor. It was available. The constraint required somebody else to say yes.
Agents are moving fastest into work where nobody thought a technical control was necessary, because a person was doing it. Research, drafting, a decision assembled across three tools and handed to a colleague who assumed somebody had already checked. The person was the control, and nobody ever wrote that down, because it never needed writing down.
So nothing was removed. There was nothing there to remove. The old arrangement ran on judgement, accountability and the plain inconvenience of a human being having to perform the action, and software dissolves the third of those before most organisations have had reason to name the first two.
The quickest available replacement is observation. Add an evaluator. Give it a policy. Log the outputs. Add a second agent to inspect the first. Put a human approval near the end for the actions that look important. The architecture feels governed, because something is watching.
Detection scales. Constraint does not. Constraint blocks actions and narrows permissions, and it forces somebody to decide in advance which things the system should simply be unable to do. Agentic systems are being bought, in significant part, to remove exactly that. The economics point at the Cricket, and they will keep pointing there.
I expect model-monitors-model to become the default oversight pattern in enterprise agent deployment. It is too useful for anything else to happen. It scales, it is cheap next to the alternative, and in most applications it will work perfectly well, which is the part that will make it hard to argue with.
Before the end of 2029 I expect a material enterprise agent incident to be traced to an AI oversight layer that certified, missed or failed to escalate another agent’s behaviour, with both systems operating broadly as designed. It will probably look more like an audit failure than an alignment story. Someone outside the loop will find it afterwards. The logs will be opened and the sequence reconstructed. The evaluator will have run. The policy will have existed. The governance slide will still be correct.
The Cricket was competent. He understood the risk, he gave the right advice, and none of it gave him any authority over the puppet holding the hammer. The Fairy already knew before she asked. Neither fact stopped anything. What stopped Pinocchio was a nose he could not get through the door.
We have known this for far longer than we have had language models. Audit knows it. Finance, security, engineering and regulation know it. Observing an action and preventing one are different functions, and we understood the difference well enough to put it in a story written for children.
Then we built systems that could act across the enterprise, and the easiest thing to place beside them was something that could watch.
The person approving that architecture can agree that independence matters, that model intelligence is not authority, that some actions should sit beyond an agent’s discretion. They can agree with every word here. And approve another Cricket.
We just preferred the story where someone else was watching.
Sources
