Long Runs area · agents that work for hours and days

Agents that work for hours. Without losing their way.

It isn't telling the agent to "work for 10 hours". It's giving it a goal with proof of done, keeping state in files and setting caps. The agent works in cycles — Codex /goal, Claude Code /goal or a loop with no interface — and the test decides when it's finished, not the agent. This area brings together the method, the execucao-longa kit and every INEMA course and project on agent loops, context and memory.

Kit guide in PT, EN and ES. Tested on real runs. All open.

01 · The method

Three pieces that let an agent work for hours without getting lost.

A long session fails in three ways: the agent doesn't know when to stop, forgets what it already did after compaction, or spends what it shouldn't. The method answers each one with a piece.

🎯 Verifiable goal

Result, constraints and verification

goal.md states the result in one sentence and lists criteria as command → expected output. No "until it looks good".

🗂️ State in files

The agent rereads, it doesn't remember

goal, plan, state, progress, failures and decisions.md hold the task; canal.md holds what compaction loses. After compacting or resuming, the agent rereads the files.

🛡️ Caps and gates

Limits on time, tokens and spend

Time, tokens and memory are capped. Credits, paid APIs, production and irreversible actions become a human gate: the agent stops and asks.

The agent picks the next useful action, tests, records and carries on. Three cycles without measurable progress = stop and call the human.

Goal→Read state→Next action→Execute→Test→Fix→Save state + commit↺
Continuity

Persistent session

The same goal and the same history over hours.

Window

Compaction

Summarizes old history to fit the window. It changes the prefix, so the cache drops right after — that's expected.

Cost

Prompt cache

Reuses the identical prefix. In a continuous session the hit rate passes 95%. Measure cached ÷ input.

02 · How to run

Two paths: you follow along, or it runs on its own.

Install once per machine (global rules, context hook and watchdog). After that, each long goal gets its own state folder and follows one of the two paths.

Interactive · you follow along

  • tools/novo-longrun.sh <project> <slug> creates the folder with the seven files.
  • Write goal.md with level-3 criteria.
  • /goal in Codex or Claude Code, with the prompt from templates/.
  • Follow without interrupting: /goals, /goal pause|resume, /side.
  • At the end, medir-sessao.py: duration, compactions, tokens and cache.

Headless · runs on its own

  • Same folder, plus prompt.md and loop.env.
  • Frozen tests: the test is the judge.
  • tools/loop-longrun.sh <folder> — each cycle is a codex exec with a time and memory cap.
  • Reverts changes to protected tests and makes a checkpoint commit.
  • Stops on its own: done, 3 cycles without progress, or the cycle cap.
# 1. state folder for the run
~/projetos/execucao-longa/tools/novo-longrun.sh ~/projetos/my-project my-goal
# 2. goal.md: result + "command → expected output" criteria
# 3a. interactive
codex   # → /goal with templates/prompt-goal-codex.md
# 3b. headless
~/projetos/execucao-longa/tools/loop-longrun.sh my-project/longrun/2026-10-01-my-goal
# 4. close and measure
python3 ~/projetos/execucao-longa/tools/medir-sessao.py <session.jsonl>

Stop and step in if: 3 cycles pass without measurable progress; the agent keeps saying "I'll finish and commit" without finishing, or reopens finished items; or it reaches the 3rd compaction in the same session — then it's a handoff and a new session.

03 · Done criteria

A good criterion: an outsider can check it, and the agent can't meet it by a shortcut.

In a long run the minimum is level 3. Below that, the agent approves itself or meets the letter without doing the work.

0 · Vague"Make the site good". Nobody knows when it's done.
1 · Subjective"Code reviewed and clean". The agent approves itself.
2 · Gameable"0 failures". You can delete a test or generate an empty page.
3 · Protected ✅"0 failures and ≥ 48 tests and tests/ untouched". Minimum accepted.
4 · IndependentLevel 3 + an outside check: real e2e, a separate evaluator, a human sample.
Five questions — each "no" drops the level. Can it be checked with a command? Is the answer yes/no? Is it impossible to meet without doing the work? Does the proof show in the output (the /goal evaluator in Claude only reads the conversation)? Does it cover function, regression and limits?

04 · Context and memory

It's not /goal that degrades. It's repeated compaction.

The goal survives compaction, but "what's done / what's left" gets lost: the agent reopens work and doesn't converge. The kit handles this before the automatic limit.

Context bands

50 · 70 · 85%

~50% → note it in canal.md. ~70% → update the state and /compact. ~85% or 3rd compaction → /session-handoff, new session and /prime. A Claude Code hook warns once per band.

Cache between cycles

Wake up before it expires

If the orchestrator sleeps between cycles: OpenAI ~30 min (wake every ~25), Claude 1 h (~55). Losing the cache costs 12.5x to 25x.

Queue mode

A growing backlog

One task per file, close only with evidence, at most 5 tasks created by the agent. Empty queue → record it and stop.

Memory across sessions

recall "term"

A SQLite FTS5 index of everything said with Codex and Claude, reindexed hourly — 90 thousand snippets, searched in hundredths of a second.

Watchdog

Flags a stalled run

tools/vigia.py every 10 min (systemd timer): a run that stalled, stopped without notice, or sits idle.

Measurement

medir-sessao.py

Reads 1.8 GB of session in ~9 s: duration, compactions, tokens, cache and tool output, per turn.

05 · Related areas

Other Events areas that connect with long runs.

06 · Projects and courses

Agent loops, context and memory: all of INEMA's work in one place.

Projects and kits

Courses (in Portuguese)

07 · Get started

The best first step: a small goal with a level-3 criterion.

Pick a task a test can prove, create the folder with novo-longrun.sh, write goal.md and run one cycle. Then measure the session.

Kit · open

Long Runs

Templates, headless loop, context hook, watchdog, measurement and recall.

Open the guide →
Kit · open

claude-session-kit

Lean Claude Code sessions: real quota, checkpoint and handoff.

Open the guide →
Sister area

Codex + Claude

One plans, the other challenges — and the goal with a stop condition.

Open the area →

Not sure where to start? Talk to us on Telegram.