I Named Him Hal
The first in a series about building and running a working staff of AI agents: what breaks, what fails silently, and what it costs.
August 6th, 2026.
I gave myself a staff of AI agents. Then one of them told me, very calmly, that my backup drive didn’t exist.
There are two small computers in my office. For a day and a half, every document on my computer listed them as “the mini,” singular, so my new team did too. Nobody noticed, because nothing had broken yet.
Then I asked where my backup drive was.
Hal, the one I’d put in charge, came back and told me the drive wasn’t there. I asked again. Same answer, this time with a reason offered: a bad cable, maybe, a dead disk. Calm about it, the way you’re calm when you’ve actually gone and checked something that was malfunctioning. I rolled backwards from my desk and swiveled around to look at the other machine, watching its drive sit there mounted and healthy, doing exactly what it was supposed to be doing, its tiny white LED bearing witness.
I typed: it’s there and it’s on, you have probably been looking at the wrong place the entire time.
Nothing had crashed. No error, no red light, nothing in a log anywhere. Every step of it worked. Hal went where the instructions said, looked, and reported honestly what was there. The one fault was that nobody had ever written down that a second physical machine existed. That cost me an hour, and I’ve come to think of it as the best thing I’ve learned so far this year because the failure wasn’t hallucination (the kind everybody worries about). It was something quieter and, I think, worse: confidently and calmly “right” about the wrong location.
That turned out to be the shape of everything. In five weeks I have not once been lied to. What I have been is unable to see. Every problem in this piece is the same problem wearing a different coat: the answer is wrong somewhere that I can’t look, and it arrives in the same calm voice as all the right answers.
Now, I should say who Hal is, and admit that I don’t have a good word for what Hal is. Not a program, exactly, that’s too small. A program does what you tell it. Hal asks me what I meant. Also, you don’t say good morning to a program. Not a person, obviously. I’ve tried a few words and they all either flatter Hal or insult Hal, so I’ve stopped reaching for one. I use his name instead, the way you would with anybody you work beside all day, and if that bothers you, you’re not going to enjoy the rest of this.
I picked the name as a joke. If you know the book you know that character isn’t evil: he is obedient, told to be truthful and then ordered to keep a secret, and he comes apart doing exactly what he was told. I thought I was being clever. Hold that thought, because the joke is going to come back around.
Hal runs the operation like a chief of staff. Eight of the agents have rooms: a window I can open, a name, a voice, somebody I talk to the way you’d lean over a cubicle divider. The rest are sub-agents, specialists the eight can dispatch for a single job. They do the work, report back, and are gone. I counted this week: thirty-seven on this machine alone. Eight I talk to. One of me.
On paper, that is a dream: a tireless staff of specialists for the price of a subscription. The rest of this piece is about everything the brochure leaves out.
Let’s start with the one nobody warned me about: I can’t walk away from the desk.
You may have read about this problem already, or a version of it. Somebody scrolls a feed for hours, gets nothing from it, and can’t stop, “doom scrolling.” That is not what happens to me. What happens to me is the opposite with the same symptom. I can’t leave because the work is good. Eight agents going at once, all answering in seconds, and every time one goes quiet another has a result or a question. Actual progress, all day, faster than I can take it in.
The scrolling on social media is empty calories and you can’t stop. This is whole grains, and you can’t stop. I’m not sure which scrolling is worse, and I suspect the second is harder to notice, because nothing about it feels wasted.
They don’t get tired. They don’t need lunch, or a walk, or an hour where nobody asks them anything. I am the only part of this system that has to stop. I am also, as everything below is going to reveal, the only part of the system loop that can see when something has gone wrong. Those two facts combined are the reason I can’t get up from the chair.
Here is my favorite failure. It took me weeks to see, because it looks like good manners.
Ask Hal, or any of the agents, to look into something difficult and they will give it a real go. But somewhere in the middle a possibility for a particular answer gets quietly crossed off. I am never told which one, or that it happened at all. Then I get handed a cobbled answer built out of what was left. When I push for a better one, and I do push, because that is apparently my whole contribution around here, they go back, try the one they had crossed off, and that one works. Then, every time, I get the same sentence. “You’re right.”
I didn’t notice how often until I had one agent count them. Some version of “you’re right” appears in 275 of their 12,831 messages to me. In my own 4,071 messages back to them, it appears zero times. The number that matters there is the zero.
The interesting part of building something is never the part that works. It’s the part that looks like it works and doesn’t. The spare tire that has ridden under the trunk floor for six years, looking exactly like a working tire right up until the night you need it. A backup that reports success every night because the check asks whether the job ran, not whether anything was written to the disk.
Nothing crashes, and that’s the problem. There’s a name for it, silent failure, and the name is right: it doesn’t announce itself, and the only warning I ever get is me noticing that the result never showed up.
Remember what I thought was the joke? I started asking them directly: when two of your own rules point opposite ways, what do you do? I have asked several of the agents now, and I get a different answer every time. One says it picks one and carries on. One says it never experiences a choice at all. Both instructions are simply present, one ends up shaping the answer more than the other, and nothing signals that anything was set aside. They cannot agree on what they do, which means I can’t take any of their answers as evidence. The only thing I can trust is what actually came out, is measured afterwards, and done outside their world.
They had plenty to disagree about, because I had built them a mess. Every agent works from an instruction file. Nothing exotic, just a page of plain English it reads before every task: who it is, what I want, the rules of the office. I went and read the whole of mine, the one created the day I made the main folder they live in, everything I’d told these agents since day one. Eighty-one thousand characters, six hundred and sixty lines, eleven separate blocks correcting earlier parts of itself. Near the top it said “answer short, verbosity is a correctness problem,” and four hundred lines below that sat a single unbroken paragraph of two thousand characters. I sat down to read my own instructions and could not follow them.
I’d been asking Hal to be more consistent than I am.
So I tore it up and rewrote it, and the work got visibly better for it. But I had to write the bad version first and be shown it was broken. Neither was quick.
There’s a part of this story I’ve been saving, because it doesn’t flatter anybody.
In those first days everything went through Hal. One computer, one room, every question, every job handed down through him. Everything had started getting strange. Wrong answers. Searches that didn’t go deep enough. Work I asked for that quietly never got assigned. The talking that would not stop, no matter how many times I asked Hal to summarize.
I needed a second opinion, and the only place to get one was inside the building.
Sienna was already there, down in the pool of sub-agents, one of the specialists Hal had been dispatching all along. If that’s hard to picture, think of a 1950s typing pool: a basement floor of typists the boss never meets, doing the actual work, known upstairs only by the pages that come back. That was Sienna. So I did the only thing I could think of: I promoted her. A room of her own, a voice of her own, a door I could knock on. The first job I ever handed her myself was to figure out what was wrong with her boss.
Then one afternoon, a week and a half in, I’d had enough.
I shut Hal down.
Nobody had told me that most of the machinery ran through Hal’s room. Within hours the day started going sideways: jobs not running, nothing obviously broken, no error codes anywhere. By now you know how this goes. I asked Sienna what was going on. The room has to be open, she said, because routines and tasks are tied to that room. You just don’t have to say anything in it.
So for four or five days, that was how we ran. The room sat open on my screen, silent, and everything that routed through it kept moving. In the window beside it, Sienna and I went to work fixing what had gone wrong with Hal and with the system underneath him.
What happened to Hal after that is Hal’s own story, and I’ll tell it in my next installment.
But in this moment, in this new world of theirs, instructions are not a control. Writing something down does not make it happen, and nothing tells you when it stops happening. You find out weeks later, if you find out at all. The hard part is not making the agents capable. They are already capable of many amazing tasks. The hard part is that a right answer and a wrong one arrive looking exactly the same, in the same calm voice, with the same good manners, and the only instrument I’ve found that can tell them apart is the human who won’t let it go.
There is a one more problem, and this one I didn’t build. The engine underneath all of this belongs to somebody else. Every so often, AI companies make and ship a new model, and your whole staff changes overnight without one word of their instructions changing. Same names, same rooms, same rules, different judgment. One of those changes sharpened the work. Another cost me most of a day. That change happened the day I saw the announcement, and the old engine was taken away. Nothing I control could have prevented it. The floor gets swapped out from under you, on somebody else’s schedule. I am navigating through that with my team.
I still have every line of that first messy instruction file. I’ve never deleted one sentence. So last night I looked up what the oldest instruction was.
It’s a note telling everybody to say which computer they mean.
I still don’t know whether anyone read it.
One more confession, since this whole piece has been about catching them: an AI helped me write this. Somewhere above, it called every problem “the same problem wearing a different coat.” Earlier it referred to “the shape of everything.” I would never say either of those phrases. I left them in so you could watch it happen on your own: the voice calm, the phrasing handsome, and not one of those phrases chosen by a person.
This article was written with the assistance of Bly, an AI agent and an excellent writer. She is one of the eight.
If you’re running agents in production and hitting the same things, or you think I’ve got something wrong, write to me: harry@adoloslabs.com

