Ryan Brooks

ralph-o-matic

Updated 2026-10-01

What it is

ralph-o-matic (ROM) is a job queue server, written in Go, that runs iterative AI coding refinement loops unattended. Point it at a git repo with a RALPH.md prompt file and it runs Claude Code through repeated iterations — against a local Ollama model, the Anthropic API, or OpenRouter, depending on the job — reading the tracking file, fixing the highest-impact issue, running the tests, committing. You queue the job, walk away, and come back to a PR.

The problem

Iterating on code until tests pass and acceptance criteria are met produces good results, but doing it by hand is tedious, and doing it against a cloud API by hand burns credits fast on work that’s mostly mechanical refinement rather than creative problem-solving. I wanted a way to hand off that grinding part of the loop — locally, for free, when a local model is good enough, or against Claude when it isn’t — without babysitting it.

How it works

The part worth knowing about is the circuit breaker (internal/executor/circuit.go). It’s a three-state machine — Closed → HalfOpen → Open — with two independent thresholds: consecutive iterations with no detected progress, and consecutive iterations that fail with the identical error message. Either threshold can be disabled by setting it to zero. “Progress” is judged from the git diff and iteration metadata, not from whether the model claims success, so a model that keeps reporting done without changing anything still trips the no-progress counter. At threshold - 1 the breaker goes HalfOpen — still running, but one more non-progressing iteration opens the circuit — and real progress from HalfOpen resets straight back to Closed. Once Open, the job stops. That’s the mechanism that keeps a stuck loop from quietly spending an entire API budget overnight.

There’s only one execution engine — internal/executor/claude.go always launches the claude binary, never a separate Ollama or OpenRouter client. What changes per job is the environment that binary sees: for Anthropic, no override, so Claude Code’s own auth routes to the Anthropic API directly; for Ollama or OpenRouter, ROM sets ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN to point Claude Code at that backend’s Anthropic-compatible endpoint instead, along with ANTHROPIC_MODEL for whichever model the job is configured to use. Claude Code itself can’t tell the difference — it’s just talking to “Anthropic’s API” at whatever URL it’s been handed. That’s also why the circuit breaker and the rest of the harness (per-iteration commits so a crash doesn’t lose work, session continuity so context carries across iterations instead of restarting cold) work identically regardless of backend: they sit outside the model call entirely. A hardware-detection step at install time (RAM, GPU, Apple Silicon) recommends a large/small Ollama model pairing — devstral on CPU with a small model on GPU, for example, if you’ve got the VRAM to split it — for the cases where the backend is local.

Status

Stable, not currently under active development. The last commit was June 26, 2026. It still builds and runs; I just haven’t had a reason to change it since.

Source on GitHub