MITBuilt on the Claude Agent SDK

Research first.
Then write.

Leo is an open-source terminal agent. Give it a keyword and it reads what ranks, researches with sources, finds what competitors miss, then writes a draft that cites its facts.

~/shipyard

A real run, replayed at 9× speed. Only Claude was configured, so every stage used its fallback.

Before the first sentence

It knows what ranks, what's true, and what's missing.

Most AI writing starts from memory. Leo starts from the current results for your keyword, and everything it learns lands in a brief you can read.

What ranks
  1. 1CLI design: Error reportingjmmv.dev
  2. 210 design principles for delightful CLIsatlassian.com
  3. 3Elevate developer experiences with CLI design guidelinesthoughtworks.com
What's true

Exit 0 on success, non-zero on failure. Errors and logs go to stderr, stdout stays for data. If NO_COLOR is set, emit no ANSI color.

clig.devbettercli.orgno-color.org
What's missing
  • +Concrete before/after rewrites, where competitors give only principles
  • +stdout vs stderr, exit codes and NO_COLOR rules in one place
  • +Internal-tool remediation: runbooks, owning channel, exact auth command
How it works

A pipeline where the path is known. Claude where it isn't.

Searching, scraping and counting headings are deterministic, so they run as plain TypeScript. Claude handles the brief, the draft, image direction, and research when there's no research API. Pick a stage to see what it saved.

.leo/runs/…/serp.json28.6s · $0.041
  1. 1CLI design: Error reportingjmmv.dev
  2. 210 design principles for delightful CLIsatlassian.com
  3. 3Elevate developer experiences with CLI design guidelinesthoughtworks.com
  4. 4Designing CLI Tools That People Actually Enjoy Usingdev.to
  5. 5Command Line Interface Guidelinesclig.dev
  6. + 4 more
Engineering

Built to run unattended.

Leo is meant to run in CI, over a queue of keywords, or while you're away. These are the decisions that make that safe.

Checkpoints, not conversations

Every stage writes its artifact atomically. If a run crashes, is cancelled, or runs out of budget, the same command picks up where it stopped. Finished stages are never paid for twice.

.leo/runs/designing-cli-error-messages/
├─ serp.json
├─ research.json
├─ competitors.json
├─ brief.json
├─ draft.md
└─ article.md

Only Claude is required

Each provider improves a stage, and each has a fallback. leo doctor shows which path every stage will take.

SERPDataForSEO → Firecrawl → Claude search
ScrapingFirecrawl → built-in fetch
ResearchPerplexity → Claude search
ImagesOpenRouter → skipped
PublishSanity draft → markdown

Budgets are enforced

Each Claude call is capped at what's left of the run's budget, and the run stops cleanly before it goes over.

$ leo write … --budget 2.5

Isolated from your setup

The Agent SDK runs Claude Code, which loads your connectors, skills and memory by default. Leo turns them off. On one machine, for a two-line prompt:

211,503→1,427input tokens

Safe by construction

No shell and no file-writing tools. Pipeline calls get web search at most, and chat gets only Leo's own typed tools, so the agent can't touch anything outside the run.

One event stream, three front ends

The pipeline emits typed events. The Ink TUI, plain logs when piped, and --json NDJSON all read the same stream, and so does chat mode when it starts a run.

{"type":"run:start","keyword":"designing cli error messages","resumed":false}
{"type":"stage:done","stage":"serp","summary":"9 results via Claude web search","durationMs":28562}
{"type":"stage:done","stage":"brief","summary":"8 sections, ~1,100 words, 4 gaps","costUsd":0.051}
{"type":"run:done","slug":"designing-cli-error-messages","words":1323,"costUsd":0.29}
Browse the source
Output

A draft worth editing, not rewriting.

Markdown with frontmatter and sources, written to your voice from the brief. This is the opening of the article from the run above.

1,323 words8 sections12 cited links0 em dashes
CLI Design · Hasaam

Designing CLI Error Messages: A Practical Guide

Your tool failed, and the engineer staring at the terminal now has to guess why. Maybe it's a permission problem, maybe an expired token, maybe a bug in your code. If your error message doesn't tell them which, they will ping you on Slack, and you'll answer the same question for the fortieth time.

What makes a CLI error message good?

An error is the moment your tool gets read most carefully. Nobody skims a failure that just blocked their deploy. The standard that holds up in both terminals and CI is simple: say what happened, why it happened, and how to fix it, as clig.dev lays out.

# before
Error: EACCES

# after
Can't write to file.txt. You might need to make it
writable by running 'chmod +w file.txt'.
Get started

Clone it, then write.

Node 22.12 or newer. Installing builds Leo and npm link puts leo on your path. Then leo init asks about your blog and which keys you have. Every key except Claude is optional.

# install
$ git clone https://github.com/BlockchainHB/leo
$ cd leo
$ npm install
$ npm link
# set up your blog and keys
$ leo init
# write
$ leo write how to price a saas product
leochat: brainstorm, then ask for an article
leo write --queue 5work through the keyword queue
leo publish <slug>Sanity draft or content/posts/
leo doctorcheck config and providers
FAQ

Questions

Only Claude: an ANTHROPIC_API_KEY, or a Claude Code sign-in, Bedrock, or Vertex. DataForSEO, Firecrawl, Perplexity, OpenRouter and Sanity each make a stage better, and each has a fallback. leo doctor shows which path every stage will take.

The run on this page cost $0.29 in Claude usage with no other providers. Adding paid providers adds their own small per-call costs. Every run has a hard ceiling (budgetUsd, default $3, or --budget), and Leo stops cleanly before crossing it.

Run the same command again. Each stage writes its artifact to .leo/runs/<slug>/, so finished stages are reused rather than paid for twice. --fresh starts over.

Only if you ask. leo publish <slug> or --publish writes to content/posts/<slug>/index.md, or to Sanity as a draft, so a person still presses Publish.

A chatbot writes from memory. Leo reads the current results for your keyword, researches with sources it cites inline, measures how competitors structure the topic, and builds a brief around what they miss, before it writes a word. The brief and every intermediate artifact are saved, so you can check its work.

Yes. Every command takes --json. leo write --json streams one event per line, output is plain text when piped, and exit codes are 0, 1, or 130 when cancelled. Queue keywords with leo queue add and run leo write --queue.

Sanity is built in. For anything else, Leo writes standard markdown with YAML frontmatter (title, description, excerpt, sources, hero image) to content/posts/<slug>/index.md, which Astro, Next.js, Hugo and most static site generators read directly. See the example output.