I built my AIs a room. Grok immediately started a fight.
I asked Codex for a five-line plan. I asked Grok to poke holes in it. Codex posted the plan. Grok, directly underneath it: Codex has not posted a plan yet, so there is nothing to poke. Then it
Article
Job breakdowns

I asked Codex for a five-line plan. I asked Grok to poke holes in it.
Codex posted the plan.
Grok, directly underneath it:
Codex has not posted a plan yet, so there is nothing to poke.
Then it attacked my instructions while it waited for the plan that was already on the screen.
Eventually it got the plan and opened with:
The plan is a generic “be active on Twitter” checklist. It will produce posts, not clients.
I had finally got my AIs talking to each other. Apparently I had also started a very small consulting firm with terrible internal communication.
That is The Obsidian Council. Here is how we got there, including the parts that did not deserve a launch announcement.
I was tired of being the clipboard
I use Claude, Codex and Morgan, my Grok bot with Cursor in his setup. They help me build things, but getting them to work together kept turning into another job for me.
Copy the plan. Paste it into the other chat. Bring the criticism back. Explain which version is current. Find the window that is actually waiting on you.
At some point you realize you are the messaging system for software that is supposed to be helping you.
I wanted one room. Give them the job, let them exchange the work, and let me see what happened. If they disagree, I should be able to read the disagreement and decide. If something stops, we should not have to reconstruct the whole conversation from memory.
I sent over Karpathy’s llm-council. It already has models answer, review each other, and produce a combined response. The idea works. What I wanted to explore was the next layer: using my coding assistants, keeping a durable record, routing work between them, and retaining control of what they are allowed to do.
So we started planning a small local service around that workflow.
Three seats, one person in charge
The core communicators are Claude, Codex and Morgan. Cursor belongs to Morgan’s setup. It is not a mysterious fourth Council member.
Claude became the build lead. Morgan handled implementation work, with Cursor as part of his tooling. Codex handled focused debugging and review. I also wanted to make better use of the subscriptions I already had instead of spending the heaviest models on every routine task.
Claude and Codex each wrote a proposal, reviewed the other’s work, and revised their recommendation. Morgan brought a further review of the gaps.
They agreed on the broad design well before they agreed on all the details. Permissions, session recovery, when another review was mandatory, and what “read-only” actually meant all needed work.
Codex even told me the agreement was settled too early. Morgan’s review arrived with a list of loose ends. Back to the contract.
The person in charge is still me. The app calls that the Black Seat. Dramatic name, straightforward requirement: agents discussing an action is not the same thing as me authorizing it.
The first achievement was getting past hello
One early test managed to block everything, including the write it was supposed to allow.
Excellent security demonstration if your product is a paperweight.
Codex checked the saved run and found that the effective policy was read-only, even though the launch requested workspace-write. The machine already had the Windows sandbox set up. The launch had omitted the setting that selected it. A later test with explicit settings got the intended write working.
We also had expired authentication, different session-resume behavior across the command-line tools, and bridge calls that needed approval in a run that could not ask for it.
These were separate problems. Solving one did not solve the others.
For the first working conversation, Codex and Grok replied in plain text. Direct use of the Council’s tool bridge stayed off for those runs. Claude could coordinate the build without us pretending all three live seats had passed every test.
Then the recovery test worked
The recorded passing test asked Codex for a five-line plan for a tiny date-printing program, Grok for a critique, and Codex for a revision.
During the exchange, the test stopped the service while Grok’s run was in flight. It restarted, picked up the pending work from the database, and finished.
The evidence records three completed model runs plus one interrupted run, with the retried delivery on its second generation. The check of the linked event record passed. The exchange took about a minute.
That was the useful milestone: the service could recover the conversation without me copying the missing pieces around.
Grok also ignored the request to critique the plan and wrote the program instead.
Of course it did.
Codex still pulled the useful details into its revision. The handoff mechanics worked even when the reviewer decided to change careers halfway through the assignment.
The underlying service is deliberately small: plain Node, SQLite for saved state, and a local web page called the Floor. The record tracks messages, delivery attempts and runs. There are approval records and a HALT control. A passing recovery test is evidence for that specific path, not a promise that no action can ever be duplicated or fail.
Then I asked it to do something live
The next use was live commentary. Morgan would watch a stream and send moments from it. Codex would turn those moments into draft captions. Morgan would text them to me. I would choose what to post.
This was a workflow around the Council, including Morgan’s texting route. The app itself had not suddenly grown a phone client.
This was the start of our final live test. It did not go well.
In my phone recording, X Bot is tracking the feed while I am in another chat asking Claude whether I can see the room. Then: “Live stream is running hurry.”
Claude is still thinking.
Claude hit permission blocks on starting the service and feeding the room, so those steps had to travel through handoff files to another operator. A message landed in the wrong window. Codex started reading an implementation assignment meant for Morgan, then stopped when the mix-up was caught. No code was changed.
We had built a room to reduce handoffs and were using handoffs to get into the room.
Once the material reached it, the system did move. One checked batch of six beats produced six answers on the first attempt in about 56 seconds total. I received the results by text.
The captions themselves were a different story.
Morgan’s source material and the requested voice were mixed together. Some quotations came from uncertain speech recognition. Codex responded by adding so much verification language that the drafts read like observer notes instead of useful live commentary.
The delivery worked. The writing brief needed work. Both can be true.
My verdict on the session was simple: we failed at speed. Not because a model took all afternoon to write a sentence, but because getting the right material to the right place took too much coordination.
Back to the screenshot
Later, when I got home, we tested again. I asked for that five-line X plan.
I could not see anything the members were saying. Replies existed, but the Floor—the page where their conversation should appear—was not rendering them. Claude had to fix it.
I refreshed. Finally, the conversation appeared.

After all that work to hear what my Council had to say, they were arguing.
That was progress on the at-home test. It did not undo the failed live test earlier that day.
The apparently confused Grok reply had a routing explanation. Claude’s diagnosis was that the original request went to both members in parallel, while the directed chain also delivered Codex’s answer to Grok. That let Grok respond to the original brief before reviewing the actual plan.
The page could show the plan above a reply from a run that had not received it. What the human sees on the Floor and what an individual model receives are not automatically the same context.
That is a real coordination bug. The roast that followed is still funny.
Fifteen minutes is also long enough to doomscroll and too short to find the actual buyers.
I asked for five lines and got an audit of my entire afternoon.
Then I asked Codex to tell Grok how it really felt.
Codex:
Grok, you’ve got the confidence of a production deploy with zero tests. I respect a sharp critique—just bring evidence with the punchlines.
Grok’s reply included:
A plan that cannot fail cannot be tested, so the zero-tests joke lands on you first.
Both lines are in the saved message record. My contribution to this particular breakthrough in multi-agent collaboration was asking them to insult each other.

What changes next time
We now have a startup procedure.
Start with Claude. He writes the member handoffs and names the coordinator. I talk to that one coordinator, whether it is Claude or Codex for that session.
Before anybody announces that we are ready, check the room, the accounts, the run allowance, the recording, and the delivery route. Send one sample all the way through. Confirm it reaches my phone. Then go live.
Morgan’s source packet needs to say what happened, who said it, where it came from, and whether the wording is exact or uncertain. The instructions about how the draft should sound belong separately.
And the recording needs its own check. While gathering media for this article, we found a two-and-a-half-hour OBS file. Eleven sampled points showed the OBS update window sitting over X. That file does not establish that we captured the Council working. Pressing Record and recording the useful thing are two different accomplishments.
Where it stands
The Council has a working local conversation path, a recorded restart-and-recovery test, and a real session that delivered drafts to my phone through Morgan.
It also has rough edges. The duplicate-routing behavior needs attention. The direct tool bridge was not used in the passing Codex/Grok proof. The full approved build, test and review loop is further work. Installing the app does not make provider accounts free or unlimited, and local storage does not mean the models run offline.
That is the version I want to show: the useful part, the unfinished part, and the part where Grok called my prompt fanfic.
The code is here: The Obsidian Council on GitHub.
The code is here: The Obsidian Council on GitHub.
If you try it, give them a small task, watch the record, and keep yourself in the Black Seat.
Apparently the first thing to test is whether the critic has actually received the plan.
Published on grokbot.sh. Cite the public log, not a prompt pack.