Result, constraints and verification
goal.md states the result in one sentence and lists criteria as command → expected output. No "until it looks good".
Long Runs area · agents that work for hours and days
It isn't telling the agent to "work for 10 hours". It's giving it a goal with proof of done, keeping state in files and setting caps. The agent works in cycles — Codex /goal, Claude Code /goal or a loop with no interface — and the test decides when it's finished, not the agent. This area brings together the method, the execucao-longa kit and every INEMA course and project on agent loops, context and memory.
Kit guide in PT, EN and ES. Tested on real runs. All open.
01 · The method
A long session fails in three ways: the agent doesn't know when to stop, forgets what it already did after compaction, or spends what it shouldn't. The method answers each one with a piece.
goal.md states the result in one sentence and lists criteria as command → expected output. No "until it looks good".
goal, plan, state, progress, failures and decisions.md hold the task; canal.md holds what compaction loses. After compacting or resuming, the agent rereads the files.
Time, tokens and memory are capped. Credits, paid APIs, production and irreversible actions become a human gate: the agent stops and asks.
The agent picks the next useful action, tests, records and carries on. Three cycles without measurable progress = stop and call the human.
The same goal and the same history over hours.
Summarizes old history to fit the window. It changes the prefix, so the cache drops right after — that's expected.
Reuses the identical prefix. In a continuous session the hit rate passes 95%. Measure cached ÷ input.
02 · How to run
Install once per machine (global rules, context hook and watchdog). After that, each long goal gets its own state folder and follows one of the two paths.
Interactive · you follow along
tools/novo-longrun.sh <project> <slug> creates the folder with the seven files.goal.md with level-3 criteria./goal in Codex or Claude Code, with the prompt from templates/./goals, /goal pause|resume, /side.medir-sessao.py: duration, compactions, tokens and cache.Headless · runs on its own
prompt.md and loop.env.tools/loop-longrun.sh <folder> — each cycle is a codex exec with a time and memory cap.# 1. state folder for the run ~/projetos/execucao-longa/tools/novo-longrun.sh ~/projetos/my-project my-goal # 2. goal.md: result + "command → expected output" criteria # 3a. interactive codex # → /goal with templates/prompt-goal-codex.md # 3b. headless ~/projetos/execucao-longa/tools/loop-longrun.sh my-project/longrun/2026-10-01-my-goal # 4. close and measure python3 ~/projetos/execucao-longa/tools/medir-sessao.py <session.jsonl>
Stop and step in if: 3 cycles pass without measurable progress; the agent keeps saying "I'll finish and commit" without finishing, or reopens finished items; or it reaches the 3rd compaction in the same session — then it's a handoff and a new session.
03 · Done criteria
In a long run the minimum is level 3. Below that, the agent approves itself or meets the letter without doing the work.
tests/ untouched". Minimum accepted./goal evaluator in Claude only reads the conversation)? Does it cover function, regression and limits?04 · Context and memory
/goal that degrades. It's repeated compaction.The goal survives compaction, but "what's done / what's left" gets lost: the agent reopens work and doesn't converge. The kit handles this before the automatic limit.
~50% → note it in canal.md. ~70% → update the state and /compact. ~85% or 3rd compaction → /session-handoff, new session and /prime. A Claude Code hook warns once per band.
If the orchestrator sleeps between cycles: OpenAI ~30 min (wake every ~25), Claude 1 h (~55). Losing the cache costs 12.5x to 25x.
One task per file, close only with evidence, at most 5 tasks created by the agent. Empty queue → record it and stop.
recall "term"A SQLite FTS5 index of everything said with Codex and Claude, reindexed hourly — 90 thousand snippets, searched in hundredths of a second.
tools/vigia.py every 10 min (systemd timer): a run that stalled, stopped without notice, or sits idle.
medir-sessao.pyReads 1.8 GB of session in ~9 s: duration, compactions, tokens, cache and tool output, per turn.
05 · Related areas
Improvement cycles in AI: the loop that learns from evidence.
Open the area → AreaA goal with a stop condition, handoff and prime — one does, the other checks.
Open the area → AreaCommanding agents that work on their own.
Open the area → AreaManaging agents: delegate, limit and supervise.
Open the area →06 · Projects and courses
Method, templates, headless loop, context hook, watchdog, session measurement and recall.
Statusline with real quota, checkpoint, memory audit and handoff before /clear.
Portable core and handoff between sessions and agents.
Open the guide → FrameworkSelf-improving loops: evidence, hypothesis, experiment, promotion and rollback.
Open the guide →Systems with AI in the loop.
Open the course → CourseLoop engineering.
Open the course → CourseFrom loops to agent graphs.
Open the course → CourseA business that learns on its own.
Open the course → CourseMastering context and tokens.
Open the course → CoursePrompt caching in Claude Code.
Open the course → CourseMemory injection via hooks.
Open the course → CourseContext engineering training.
Open the course → CourseOne does, the other checks.
Open the course → CourseDelegate, verify, iterate and scale with Codex.
Open the course → CourseDevelopment with agents: plans, tests and verification.
Open the course →07 · Get started
Pick a task a test can prove, create the folder with novo-longrun.sh, write goal.md and run one cycle. Then measure the session.
Templates, headless loop, context hook, watchdog, measurement and recall.
Lean Claude Code sessions: real quota, checkpoint and handoff.
Open the guide →One plans, the other challenges — and the goal with a stop condition.
Open the area →Not sure where to start? Talk to us on Telegram.