Evaluation produces lines, not impressions
One line per failure in the log: date, what broke, the smallest possible fix, and whether it was an instruction or an infrastructure problem. After ten lines, the pattern shows up.
Cultivated AI area · it isn't programmed, it's cultivated
Nobody writes, line by line, a large model's ability to draw analogies, write code or plan a company. The labs create the conditions — architecture, data, objective, compute, training, feedback — and the capabilities emerge. This area takes that idea and shows how to apply it at the level where you actually work: the agent. Eight elements, one cycle, three gardens — your personal life, your Jarvis and your business — and the set of files that makes cultivation happen.
The model is the same one your competitor has. The garden isn't.
01 · The idea
You choose the seed, the soil, the water and the light. You don't draw each leaf. And here is the point almost everyone blurs: there are two levels of cultivation, and only one of them is yours.
Level 1 · the lab cultivates the MODEL
architecture + data + objective + compute + training + feedback → emergent behavior
It happens at Anthropic, at OpenAI, at Google, at DeepSeek. It costs billions and you take no part in it.
Level 2 · you cultivate the AGENT
role + context + tools + rules + examples + memory + evaluation + feedback
The model arrives finished, with its weights frozen. What you cultivate is the envelope around it — and that envelope is what turns the very same model into a confused intern or a dependable professional.
The honest consequence. In use, the model does not "learn" on its own. What learns is the system: your context file, your memory, your rules, your examples. When you say "my AI got better", what got better was the garden, not the seed. And that's great news, because the garden is yours.
Traditional software
Cultivated AI
Every agent, whether it's your personal Jarvis or a sales agent inside a company, is cultivated with the same eight ingredients. The first four make the agent work. The last four make the agent improve — and most people stop at the fourth and then complain that "AI doesn't learn".
What is it accountable for?
One thing, well defined. One role at a time, not "does everything".
What does it need to know?
About you, the business, the customers, the processes. One good page is worth more than twenty dumped in.
What can it operate?
Calendar, CRM, email, browser, files, APIs, MCPs. Start with read-only.
How far does it go alone?
What it does without asking, what needs confirmation, what it never does.
How does a good professional do it?
Real annotated cases — including the bad ones and the approved exceptions.
What carries over between sessions?
Preferences you discovered, decisions, what has already been tried. A short index, not an endless diary.
Is it good? Compared to what?
Quality, cost, time, errors and outcome. A weekly grade or a fortnightly scorecard.
What do you do with the evaluation?
Turn every failure into a change of context, rule, example or tool. In the file, not in the chat.
Read this before getting carried away. Cultivation is not agricultural magic.
Where the idea comes from. The phrasing "grown, not built" appears in Dario Amodei's essay on interpretability, which credits it to Chris Olah. The project's research report gathers the sources — and also records what it was not possible to confirm.
Read the research report (in Portuguese) →02 · The cycle
What changes is what you wrote down about where it went wrong and what you adjusted in the environment. That's why the cycle is the heart of the method: without it, the first four elements add up to an agent that stalls.
One line per failure in the log: date, what broke, the smallest possible fix, and whether it was an instruction or an infrastructure problem. After ten lines, the pattern shows up.
If you've explained the same thing three times, the model isn't being stubborn: a rule is missing. Feedback only counts when it goes back into the context, the rules, the examples or the tools.
Level 1 proposes and you execute; level 2 executes and you review; level 3 executes and reports. A task only moves up after weeks with no entry in the failure log. It never starts at level 3.
03 · Garden 1 · personal life
Most people open the chat, ask, copy the answer and close it. Every conversation starts from zero: it's like hiring a brilliant consultant and wiping their memory every morning. In your personal life cultivation is cheap — three or four text files and one weekly habit. No CRM, no API required.
Role: an advisor who doesn't decide. Context: your criteria and past decisions with their outcomes. Rule: always present the opposing option. Emergent: it starts reminding you of your own patterns.
Role: coach. Context: your real routine, constraints, what you've already dropped. Rule: don't prescribe — suggest and tell you to ask your doctor. Evaluation: weekly adherence, not motivation.
Role: spending analyst. Tool: a spreadsheet exported from your bank. Rule: never move money, only show it. Emergent: spending patterns you had never seen.
Role: tutor. Context: what you already know and how you learn best. Evaluation: it asks, you answer, it notes where you get stuck. The memory file becomes the map of your gaps.
Role: an editor with your voice. Examples: five of your own texts. Rule: don't change the tone, only the clarity. Emergent: after a month, the first draft comes out almost finished.
1) Read the memory file. 2) Write three lines in the failure log. 3) Make the fix in the file, not in the chat. 4) Delete from memory whatever is no longer true. Eight weeks of that and you have an AI that looks like nobody else's.
04 · Garden 2 · your Jarvis
"Jarvis" is the popular name for the agentic personal assistant: an AI that doesn't just answer but acts in your environment — reads and writes files, touches the calendar, sends messages, browses, runs commands, remembers what happened yesterday. Two Jarvises on the same model can be a disaster and a dependable partner. The entire difference lies in the lines below, which are text files you write and revise.
What comes from the lab
MODEL (fixed, frozen weights)
The same for you and for your competitor.
What you cultivate
context · memory · tools · rules · skills · evaluation + feedback
This is context engineering: deciding what enters the window on each task. Six lines in a file, not six months of engineering.
It drafts the email, you send it. Every agent and every new task starts here.
It organizes the folder, you check the result. Only after a few weeks with no entry in the failure log.
It runs the daily routine and sends you the summary. Reserved for tasks with a low cost of error and a clean track record.
After two months, that Jarvis ships an entire project from a one-line request.
Not because the model got better. Because the garden was ready.
Personal agents have access to your life. The 2026 incidents with exposed OpenClaw instances showed the pattern: thousands of agents open on the internet with no password, credential files leaking, malicious emails instructing the agent to hand over session cookies. None of that is the model's fault. It's a garden with no fence.
Minimum rules
What went wrong out there
Where this is documented. The cases, the CVEs and the scans are listed with sources in the project's research report, along with what the research could not confirm. Nothing here is an estimate of ours.
05 · Garden 3 · business
You don't program a salesperson line by line: you give them a role, context, objectives, rules, tools, examples and feedback. With agents it's the same. And when projects fail, the reason is rarely the model — it's a lack of context, process and feedback. MIT's report on the generative AI divide traced the root cause of pilot failure to organizational factors, not technical ones.
Rigid automation
"If A happens, do B, then C."
It breaks on the first case nobody anticipated.
Cultivated agent
"Your role is to qualify leads. Here are our criteria, our CRM, examples of good and bad leads, your limits and the expected outcome. Execute, log what you did and learn from the evaluation."
The agent doesn't get instructions. It gets a working environment.
The first agent always at the "it proposes, a human executes" level. One process, one agent, one metric.
It scores the opportunity, recommends the next action and drafts the follow-up. It logs what it did.
Based on the knowledge base, with rule-driven escalation. The company's best-documented process is usually where the agent flourishes first.
Entries classified and friendly collection messages drafted — never sent without confirmation.
Draft content and a weekly metrics report.
Reading documents, extracting data and checking against a checklist.
Résumés screened by written criteria and answers to internal policy questions.
It replaced hundreds of support agents with AI in 2024, admitted a drop in quality in 2025 and went back to hiring humans for complex cases.
Using AI became a baseline expectation: before asking for a new hire, the team has to show why AI can't handle it.
It reports hundreds of millions in recurring revenue from agents in customer support.
It announced going "AI-first", faced public backlash and walked it back.
HBR research shows that nearly half of workers receive AI-generated content that looks good and is useless, costing hours of rework.
It's not "prompt engineer". It's whoever writes the specs, curates context, maintains examples, reads scorecards and runs the feedback ritual. The manager doesn't have to be better than the AI at the task; they have to know how to create the environment in which the AI produces the right result.
Competitive advantage won't belong to whoever has the best model.
It will belong to whoever has the better context, processes, tools, feedback and agent management.
06 · Practical kit
The five files are the same in the personal garden and in the company one — the content changes, the structure doesn't. The project ships these files ready to use, with a fictional example filled in, in three downloadable kits.
Role, expected outcome, human owner, what it does alone, what needs confirmation, what it never does. With a version number and a date.
"About me" or "About the company": who it is, objectives, constraints, how you like to work, active projects, decisions that don't get reopened.
Good, bad and approved exception — with the reason why. Twenty cases, including the ugly ones. Examples of easy cases only produce an agent that only handles easy cases.
One line per failure, most recent at the top: date, what broke, the smallest possible fix, instruction or infrastructure. No narrative.
A human-reviewed sample, average cost per task, time, serious errors, business outcome, and the decision: hold the level, move up or step back.
Evaluation produces lines and numbers. Feedback turns each line into a change of context, rule, example or tool. The spec goes from v1.0 to v1.1 and what changed is on the record. None of this changes the model. All of it changes the agent.
The five files already structured, with a fictional example filled in, in three versions. Zips ready to go.
Open the kits → GeneratorFill in the spec, context, examples, scorecard and failure log right in the browser. A static page — nothing leaves your computer.
Open the generator → AssessmentEight questions — one per element —, a score, where to start and which kit to use.
Take the assessment → Skills/cultivar and /revisao-semanalFor Claude Code and Codex, with install.sh. One builds the garden, the other runs the ritual that closes the cycle.
A complete lead qualifier (system prompt and n8n flow), plus support triage, financial reconciliation and a marketing report. These are templates — they have not been run against a real CRM or a real n8n.
See the packs on GitHub → Full textConcept, personal life, Jarvis, business and the practical kit — to read, copy and adapt.
Open the chapters (in Portuguese) →07 · The course and what comes with it
The Cultivated AI course is a single page holding the landing and the whole content. The repository is a GitHub template: "Use this template" creates your own copy with the kits, generator, assessment, skills and packs inside.
The concept, the eight elements, the cycle and the three gardens on a single page. It's the starting point of this area.
Open the course → RepositoryChapters, kits, generator, assessment, skills, packs, sources and research. A template ready to copy.
Open the repository → ResearchThe origin of the concept, recursive self-improvement, corporate adoption, personal agents and criticism — with whatever could not be confirmed marked as such.
Read the research (in Portuguese) → Founding textThe first of the two texts the project grew out of.
Read the text (in Portuguese) → Founding textThe second text: the same thesis applied to the organization.
Read the text (in Portuguese) →The course and the tools are also available in Portuguese and in Spanish, with translated kits, generator and assessment.
If you're on the side of a company that already has agents and needs to manage them — roles, authority limits, cost per outcome and evaluation —, the AI Management area continues this conversation. And for building the agent itself, the AGI-ready area.
Open AI Management (in Portuguese) →08 · Start now
You don't need a big project. You need one well-defined role, one page of context and the habit of writing down where it went wrong. The course and the tools are open. What changes when you join the ecosystem is having people, material and support instead of doing it alone.
The single page with the concept, the eight elements, the cycle and the three gardens — plus the tools that come with it.
/cultivar and /revisao-semanal skillsThe community of people applying AI to real businesses: a feed curated by Nei with what matters, groups by topic and the supporting material from every training.
The platform for anyone who wants to use AI to grow in practice: the ecosystem's projects, the training programs and the Brain, with the content organized for reference.
Not sure where to start? Talk to us on Telegram.