Skip to content
Bot jobsJob breakdowns

How I Hand Opus 5.5 More of the Work: 5 Prompts for a Model That’s Less Dumb

Models aren’t necessarily getting smarter. They’re getting less dumb, and being smarter and being less dumb are two very different things. Here’s what changed in how I work and the setup I use: Claude

Cole HollanderImported from X9 min read
Colehollander10x article
See this runHouse 479 · 00663

Article

Job breakdowns

Models aren’t necessarily getting smarter. They’re getting less dumb, and being smarter and being less dumb are two very different things.

Here’s what changed in how I work and the setup I use: Claude Code, T3 Code, and Grok Bot. I’ve included five prompts for handing over more of the process, checking the work, and preparing for tomorrow.

1. Less dumb beats smarter

Recently, I asked Opus 5.5 to look at the Support Agent at a very high level and find concrete ways to improve it. That could be as small as bug fixes or as big as completely changing the structure of how it works.

It went off for about 15 minutes. I think it used a few GPT-6.1 Sol reviewers, which worked out really well. It found bugs, cut out a bunch of code, and created pull requests. The Support Agent ended up running a lot faster.

I said, “Okay, go ahead and merge those,” and it did.

It did such a good job of getting it right the first time, which didn’t use to be the case.

The less a model makes really stupid mistakes, the more I’m going to want to use it, and the more of the process I can hand over to it.

I think back to older Anthropic and OpenAI models. They could get work done, but would still make stupid decisions sometimes. I couldn’t really leave one to go do things for hours without it getting stuck or making mistakes.

Finding a model that you trust is so important. Opus 5.5 does such a good job getting it right the first time and working on its own.

It feels weird relying less on skills and custom hooks, but the model doesn’t need those things as much as previous ones did.

It still makes mistakes. But the “crash-out” factor is a lot lower.

Theo also described giving Opus 5.5 a rough idea and getting back a pull request with a video demo of the changes. His post

2. What I’m building with it

A couple of the things I’ve been working on with Opus 5.5 are Otto Rabbit and updates to the Support Agent.

Otto Rabbit is a Claude Code skill I’ve been building with AI.

It’s designed to run as a Claude Code routine in the cloud, spin up a copy of the Order Desk codebase, and review PRs and comment on them. It’s doing the kind of work CodeRabbit does, but with Opus 5.5.

What’s really cool is the idea of it improving from developer feedback. If it sees a developer catch something it missed, it can learn to watch for that type of problem on the next PR.

The Support Agent is a custom Claude Code plugin. You can put in a ticket, and it reviews the ticket and searches across the different sources.

Opus 5.5 has been showing really sound judgment. I’ve been impressed with the work it’s getting done.

3. My setup

I mostly use a local Claude Code session when I’m building because I have more control over it and it can see my local files.

A few reasons to choose a local session:

  • Work with local files and unfinished changes.

  • Use local tools and your development environment.

  • Work through requirements that are still changing.

  • Review changes and give feedback as the agent works.

Right now, I’m using T3 Code as the harness. It does a really good job with the things I want out of a harness.

It supports multiple providers, so Claude Code can delegate to Codex and other harnesses natively. I can go into the thread and see what those agents are doing.

I use Opus 5.5 for the main work and GPT-6.1 Sol as a reviewer.

Anthony Kroeger also described using Opus to code and Sol to review. His post

There are a bunch of other things I like about T3 Code:

  • The sidebar shows which agents are working, waiting for me, or finished.

  • I can see which provider each session uses.

  • Threads are easy to rename and organize.

  • It has a built-in browser and handles visualizations well.

  • It’s customizable, gets frequent updates, and includes usage insights.

The sidebar is the best part for me. It does a really good job of showing me what’s happening.

4. How to hand over more

If you’re using Opus 5.5 and you don’t notice much difference from the previous models, I don’t think you’re being ambitious enough with how much of the process you’re handing over.

When I give it a few minutes of me just voice prompting it, it does a really good job of inferring what I want and getting the work done. The more specific I am, the better it is.

When I’m not specific enough, it can go off and do something I didn’t ask it to do. Sometimes I just haven’t given it enough information.

What I include in that first prompt:

  • Goal: What are you trying to accomplish?

  • Context: Where can it find the relevant information?

  • Constraints: What should it preserve or avoid changing?

  • Finished: What result do you want back?

  • Review: Where should it stop for you?

I decide what it can complete and what it should bring back to me. Here are four prompts for the work itself. The fifth, for preparing for tomorrow, is in the personal-agent section.

Prompt 1: Fix a reported issue

Use this when you have a specific issue and want the agent to carry the fix through the checks.

Prompt 2: Review the implementation

Use this after the work is done, when you want it to review and fix what it finds.

Prompt 3: Own the issue through completion

Use this when you want to hand over the investigation, implementation, and review together.

Prompt 4: Move a project forward

Use this when the project goals are clear, but you want the agent to choose the next useful work.

One of my favorite prompts came from someone else. This is my adaptation:

5. Context and data safety

I store my work and updates in Notion. If my bot wants to know what I’ve been working on or what the latest updates are, it can look there.

With Slack, it can see that I was talking to someone yesterday about something and forgot to follow up. Or that someone was going to send me something and hasn’t, so maybe I should message them about it.

I still write all my Slack posts myself. The agents help me notice the things I should follow up on.

The more context an agent has, the more it can help me. Knowing what’s on my calendar, in my email, or in the codebase matters a lot.

Lauren from the Grok Bot team gave an example of combining a Slack channel about new features with CRM information to work out which customers to contact. Her post

That also means paying attention to customer data. I’ve been drafting reports from Help Scout analytics, and I’m trying to make sure there’s good masking in place at every level when I run them.

The more context you give AI, the more useful it is, but the more context it has about you. I want to stay aware of the issues that come with that.

6. Personal agents

For building, I still lean on local Claude Code. Grok Bot has been useful for knowledge work, especially when I’m on my phone or on the go.

Grok Bot

When I first tried Grok Bot, I loved the design. I thought, “This is a compelling idea!”

But there weren’t very many things I could connect it to, and it was annoying having to sign in when I wanted it to access something on the computer. It was also really slow, and the model wasn’t that good.

The speed and connection problems have improved drastically. It responds really quickly now, and the work is higher quality.

You have one primary bot doing most of the work, and it can delegate to specialist bots for certain workflows. Most of the time, I only need to deal with one thread. Grok Bot documentation

You have one primary bot doing most of the work, and it can delegate to specialist bots for certain workflows. Most of the time, I only need to deal with one thread. Grok Bot documentation

  • Native X search is useful for me because I’m on X all the time. Being able to send it posts and have it work with them is convenient. Official announcement
  • The team has announced support for models from other providers, including Opus 5.5. Model announcement

My bot now has its own email address, which is really useful.

Recently, I told it I was off for the day and asked it to finish up what it could. I didn’t know what I had left, and I wanted it to figure out what was important so I’d be prepared when I came back.

I also told it that some reports needed manual time it couldn’t put in.

It got three or four things done, and when I came in the next morning, it had a great list for me.

A lot of my work still needs my own time and manual effort. I’m still learning how to delegate work to it.

Dots

Dots is OpenAI’s version of a personal agent. It has its own cloud computer, can use connected apps, and can keep working while you’re away. Introducing Dots

Dots is OpenAI’s version of a personal agent. It has its own cloud computer, can use connected apps, and can keep working while you’re away. Introducing Dots

I named my Dot Margot. The OpenAI team has shipped a lot of improvements, but I still find myself reaching for Grok Bot when I want a personal assistant.

The great part about Dots is that it has the same personal context as your ChatGPT account: plugins, memories, chats, and so on. It can also delegate to other Codex threads.

If I spend some time learning how to use it well, it could end up being a powerful delegator.

The call feature is pretty bad in my experience. It glitches out, and overall the product feels half-baked and needs a lot of work.

Outside of simple daily briefs, I’m not using it much right now. I think it probably will be good soon. For now, though, it needs more work.

Where a personal agent fits

Some tasks to try:

  • Research across different sources.

  • Preparing for the day.

  • Finding unfinished tasks and missed follow-ups.

  • Work that depends on Notion, Slack, email, or calendar context.

Grok Bot can also hand tasks off to cloud coding agents, so there’s some overlap.

I’m still in the phase where I trust my agents to do the middle of the work more than the beginning or the end.

The beginning is the idea, thinking about it, and brainstorming. I might hand some of that over, but most of it I’m still doing myself. From there, I’ll give it the go-ahead to do a lot of the work.

At the end, when it’s time to respond to someone, give an update, or do the final touches, I come back in.

Start small

Give your agent one recurring task, and as it builds trust, give it more.

  1. Choose a recurring task. Preparing for tomorrow, reviewing project updates, or finding follow-ups are useful starting points.

  2. Give it relevant context. Point it toward the specific projects and conversations it needs.

  3. Explain what finished looks like. Ask for a result you can inspect.

  4. Choose the review point. Decide what it can complete and what it should bring back to you.

  5. Check the actual work. Read the sources, inspect the changes, and see whether it missed anything.

  6. Expand the assignment as it earns trust.

Prompt 5: Prepare me for tomorrow

Use this when you want a list of unfinished work and follow-ups to review before the next day.

I think personal agents could help the team at Order Desk get more done, and it’s worth trying. The hard part is learning to direct an agent and check its work, just like a teammate.

I don’t think one company will win forever, so I want us to stay flexible about tools while keeping our work context in Notion and Slack.

My stack

  • Claude Code: Local building with Opus 5.5, with GPT-6.1 Sol reviewing the work.

  • T3 Code: The harness I use to manage the sessions and see what the agents are doing.

Grok Bot: My personal assistant for knowledge work and follow-ups. Documentation

  • Grok Bot: My personal assistant for knowledge work and follow-ups. Documentation

Dots: Still experimenting, mostly with daily briefs. Introducing Dots

  • Dots: Still experimenting, mostly with daily briefs. Introducing Dots

Bookmark this for the prompts, and follow me for more of what I’m building with agents.

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu