llmthing Get in touch

Find the session · own the loop

Agents do the work. The runtime owns the loop.

Coding agents leave sessions everywhere and decide for themselves when they're done. llmthing is two small tools for working with them: OSIRIS finds and resumes any session on the machine, and PHARAOH runs the build, verify and review loop, so checks decide instead of the conversation.

Act I

Working with agents today.

Several agents, several worktrees, several terminals. Two problems show up every day: finding the work, and trusting that it's done.

01 · Where was it?

"Where was that session?"

You remember the work, not where it lives: which agent, which repository, which worktree, which terminal. So you grep through session stores.

~
$ ls ~/.claude/projects ~/.codex/sessions | wc -l
214
$ grep -rl "migration" ~/.claude/projects | head -3
…/-home-you-src-api/5f2c….jsonl
…/-home-you-src-api-wt-3/91ab….jsonl
# two worktrees. neither says which branch it was on.

02 · Is it done?

"Stop when it works."

So it stops when it thinks it works. Whether verification was enough is a judgement call made inside the conversation.

assistant
Build this, test it, fix whatever fails, review it, and stop when it works.
Done! I've implemented the feature and all tests pass. ✓
$ cargo test
test result: FAILED. 146 passed; 2 failed

03 · Round and round

Then it tries again. And again.

No retry limit and no point where a human is asked. Every decision about the loop lives in a chat that forgets its own start.

assistant
Two tests fail.
You're right, let me fix that. Tests should pass now.
One still fails.
Let me try a different approach…
Now the first one fails again.

That was one task.

commands
3
turns in the chat
4
false dones
2
decisions the checks made
0

Here's the same work, with llmthing.

Act II

The llmthing way.

Two small tools on top of herdr, the runtime that already runs the agents. OSIRIS first, because it's smaller and useful on its own.

01 · OSIRIS

Find any session, by what it was about.

Every agent's sessions on the machine, indexed with their repository, worktree and branch. Search for what you remember and get it back.

~
$ osiris postgres migration
search: postgres migration
api-server Claude 2 days ago
Fix migration deadlock
~/src/api fix/migration-lock
billing Codex 5 days ago
Backfill invoices table
~/src/billing-wt-2 feat/invoice-backfill

02 · Resume

And resume it where it lived.

OSIRIS brings you back into the right directory, on the right branch, with the right agent, however long ago it was.

~
$ osiris resume 1
✓ resumed in ~/src/api on fix/migration-lock · Claude

03 · PHARAOH

The loop, written down.

Build, verify, review and accept are stages in a file. Verification is commands with exit codes, and retries have a limit and carry the failure along.

workflow.yaml lines 1–25
build:
agent: codex
verify:
run:
- cargo fmt --check
- cargo clippy
- cargo test
on_verify_failure:
retry: build
max: 3
include:
- verify.output
review:
agent: claude
on_blocked:
human: true
accept:
when:
- verify.passed
- review.approved

04 · A run

The runtime decides, not the chat.

The agent builds; cargo decides whether it passed; one retry goes back with the output; the reviewer's verdict lands in a report. Blocked or out of retries, it stops and asks you.

~/src/api
$ pharaoh run workflow.yaml
BUILD codex implemented the task · 3 commits
VERIFY cargo 2 of 148 tests failed
BUILD codex retry 1 of 3 · with verify.output
VERIFY cargo 148 of 148 passed · clippy clean
REVIEW claude approved · 1 note left in the report
ACCEPT verify.passed · review.approved
DONE report and branch handed off

Act III

The rules that make it a pipeline.

Without these it's an agent chat room with extra steps.

  1. 1

    The gate is a command

    A stage passes when its command exits zero. The reviewer is advisory, and its verdict goes in the report, not in the decision.

  2. 2

    The handoff is a commit

    Each stage hands on a branch and a written report, never a transcript, so the next stage doesn't need the last one's context.

  3. 3

    A human, on purpose

    Blocked, ambiguous or out of retries means the run stops and asks. And every run is checkpointed, so it survives the machine restarting.

Working with agents, both ways
Question Today llmthing
Where was that session? grep the session stores osiris
Is it done? the agent says so the checks pass
How many retries? until you notice a limit, with the failure attached
When to ask a human whenever it happens to when the workflow says

Running more agents than you can keep track of?

Tell me how you work with them. OSIRIS is next on the bench.

Get in touch