JEV: AI decisions
in practice.

A place to explore the project, study the course, and bring structured decisions to your code. Try the 20 cases, go through the 36 lessons, and discover the 17 packages by area.

Public lab without an API key, with simulated responses. HTML v2 course in Portuguese, with progress tracking, questions and notes.

Banner of the initial edition, with ten packages. The current collection includes 17 packages.

From context to decision. From decision to evaluation.

Jev is TypeSafe's structured decision model. The Jev Decision Lab is INEMA's educational project to formulate questions, compare answers and understand when a decision needs review.

The model receives context and criteria; your system remains responsible for the actions. The lab separates this choice from text generation and tool execution.

Consult TypeSafe's official documentation

Three ways to ask

  • Choice: choose among explicit alternatives.
  • Noul: estimate the probability of an affirmative answer.
  • Score: evaluate a rubric with ordered levels.

What you can already use

  • Search among 20 original cases and combined questions.
  • Textual or JSON context editor and decision criteria.
  • Import requests and export results.
  • Didactic policy of thresholds, abstention and human review.
  • Cost estimate with fallback, review and infrastructure.
  • Batch experiments: rules, direct Jev, hybrid and replay.
  • Metric comparison and report viewer in the browser.
  • Python client, local server and integration via OpenRouter or TypeSafe.

20 situations to try out

Each case includes context, questions, criteria, simulated answer, explanation and next step. The links open the corresponding case in the lab in Portuguese.

Agent routing

Dispatch to a specialty without expanding permissions.

Comment quality

Separate correctness and usefulness before proposing a revision.

Log priority

Use an operational rubric, without inferring the entire incident.

The course: Jev in practice

Three tracks, 12 modules and 36 lessons with theory, examples, exercises and answers. Open each module to see the lessons and access the full material.

The conceptual journey covers modules 1–8. Modules 9–12 use JSON, terminal and Python. The course estimate is 18 hours, including exercises and project.

HTML v2 course available in Portuguese. The links below open lessons with progress tracking, questions, notes and exercises. A personal worksheet follows one decision from the beginning to a pilot.

Understand

Module 1 — Where Jev fits

Choose between a rule, a bounded decision and a generative response.

  1. Decide, generate or calculate
  2. Context, question and options
  3. Read promises carefully
Read the three lessons of the module
Module 2 — The three types of question

Distinguish Choice, Noul and Score based on the required response type.

  1. Choice: select an alternative
  2. Noul: probability of yes
  3. Score: ordered levels
Read the three lessons of the module
Module 3 — Confidence and error

Use uncertainty without turning it into automatic authorization.

  1. Probability and confidence
  2. High confidence, wrong answer
  3. When to request a review
Read the three lessons of the module
Module 4 — Cost and choice of the first case

Calculate viability including work that happens after the model.

  1. The ten‑thousand‑token account
  2. Cost of the entire process
  3. Choose a pilot
Read the three lessons of the module

Apply

Module 5 — Service and transcripts

Apply decisions to messages and distinguish lack of evidence from a negative response.

  1. Forward tickets
  2. Sales and scheduling together
  3. Missing information
Read the three lessons of the module
Module 6 — Documents and evidence

Assess what a passage supports without extrapolating to external truth.

  1. Is a statement supported?
  2. Clause checklist
  3. What changed in the conversation?
Read the three lessons of the module
Module 7 — Models and agents

Routing, verification and execution as separate responsibilities.

  1. Choose a model
  2. Verify a step
  3. Choose an agent
Read the three lessons of the module
Module 8 — Browser and sensitive applications

Distinguish candidate selection from execution and professional supervision.

  1. Select element by text
  2. Organize a fictional clinical review
  3. Organize fictional alerts
Read the three lessons of the module

Build and evaluate

Module 9 — Integrate without mixing responsibilities

Build a request and handle errors while keeping credentials on the server.

  1. Prepare the state
  2. Understand a request
  3. Handle operational failures
Read the three lessons of the module
Module 10 — Measure truth quality

Compare methods without confusing tuned examples with independent evidence.

  1. Build human reference
  2. Compare alternatives
  3. Choose and freeze thresholds
Read the three lessons of the module
Module 11 — Operate and observe

Prepare logs, monitoring and feedback without losing control of the operation.

  1. Observe before automating
  2. Record and understand failures
  3. Update or roll back
Read the three lessons of the module
Module 12 — Final project

Deliver a well‑grounded adoption decision, including when the answer is not to automate.

  1. Specify the screening
  2. Evaluate the proposal
  3. Decide adoption
Read the three lessons of the module

12 labs, materials and final project

Fictional, author‑created data with activities and answer keys for self‑grading. Solve first and check the answers afterward.

L6 — Browser

Five candidate lists to choose, abstain, or request review.

L10 — Code

Six situations to separate correctness, usefulness and test evidence.

Your final project

Specify the problem, taxonomy, questions, policy, evaluation and economics. Conclude whether there is evidence to adopt, collect more data, or not automate. The rubric covers formulation, policy, evidence, economics and reproducibility.

Open the model and the final project rubric

17 packages to apply at work

Context, criteria, request, fixture and instructions in each package. All share the executor and the Python core; you can start offline and later set up a real query.

Customer support

Classify tickets to the correct queue.

Insert the evaluation between ticket creation and queue selection.

Sales and CRM

Qualify opportunities according to explicit criteria.

Trigger after receiving a contact form. Add your catalog and commercial criteria to the context.

E‑commerce

Organize post‑sale requests.

Evaluate the ticket with only the strictly necessary order data.

Marketing and content

Check whether a text follows the provided brief.

Evaluate the draft before the editorial approval stage.

Software development

Triage bug reports for the responsible component.

Use the issue title, description and minimized logs as input.

Agents and automations

Choose among registered skills without executing them.

Place the classifier before the tools dispatcher. Use only IDs from the allowed catalog.

Education and training

Apply an explicit rubric to support the teacher’s review.

Use fictional exercises first; then evaluate minimized responses against the teacher’s rubrics.

Document Management

Check mandatory information in documents.

Extract text before this step and attach the checklist.

Operations and Projects

Classify pending items for the responsible person's review.

Use project updates as input before the follow‑up meeting.

Internal research and knowledge

Assess whether a provided excerpt supports a claim.

Use after retrieving excerpts from the knowledge base and before drafting a response.

Inbox

Separate messages and flag those that need review.

After receiving the email in your backend; send only the required body and profile.

YouTube comments

Prioritize questions and content suggestions.

After importing comments via your application's authorized integration.

Communities and courses

Locate help requests and expressed dissatisfaction.

After a post or question is entered into the learning environment.

Meeting quality

Check decision, next step, owner, and deadline.

After the meeting has been authorizedly transcribed.

Clip selection

Assess whether a transcript excerpt supports an independent clip.

After transcribing and segmenting in inemavox; send the text and timestamps, not the video file.

Organization of notes and audio

Separate tasks, ideas, journal entries, and references after transcription.

After a textual note or a transcription in your bot; keep the original in the source system.

Feed curation

Prioritize textual relevance according to declared interests.

After importing posts through an authorized mechanism.

The packages are in the repository: there is no PyPI distribution yet, no universal installer, and no ready‑made connector for n8n.

Start simple. Move forward when you need to.

In the browser

Open the lab, choose a case and examine the instructional answer. Edit the questions, import a request, export the result or open a report.

Open the public lab

The public site does not call an API and does not request credentials. Changing a case invalidates your previous simulated answer.

On your computer

With Python 3.10+ and Git, clone the project and start the local server. The core uses the standard library.

git clone https://github.com/inematds/jev.git
cd jev
python3 -m jev_lab serve

Open http://127.0.0.1:8765 in the browser.

From offline package to real inference

python3 -m pacotes.executar --list
python3 -m pacotes.executar atendimento
python3 -m pacotes.executar atendimento --live --provider openrouter

Without --live, the package uses a simulated fixture. Real mode requires a key configured in the backend and may consume credits. The guide shows how to set up OpenRouter and TypeSafe while keeping credentials out of the browser.

Batches, resumption, and evaluation by question

Validate a JSONL file without calling the API. With --live, process events with limited concurrency and resume saved results. Compare Choice, Noul, and Score with explicit references; the new demos remain simulated.

python3 -m pacotes.executar reunioes
python3 -m pacotes.qualidade reunioes
python3 -m pacotes.lote reunioes data/reunioes-eventos.jsonl
See practical flows and executor limits

In Codex, Claude Code, and OpenPCBot

Use the jev-decidir skill in Codex and Claude Code. In OpenPCBot v3, /jev observar compares route suggestions; /jev shows the result and /ajuda jev explains the limits. The bot maintains its own gateway and does not execute Jev's suggestions.

Install the skill for Codex and Claude Code

A gateway for your system

When a Jev query originates from multiple points in the same system, it makes sense to funnel everything through a single gateway. jev-gw does that: daily spend limit checked before each query, cache for identical requests, cost and latency logging per call, and a conservative failure mode — Jev down, wrong key or limit exceeded return a human review instead of breaking the caller. Python library, HTTP service and CLI, all using the standard library.

from jev_gw import decidir

saida = decidir(pedido)
if saida['acao'] == 'suggest':
    encaminhar(saida['resposta']['answers']['fila']['choice'])
else:
    fila_de_revisao(saida['motivo'])

Measured on 09/22/2026: a real query took 610 ms for US$ 0.0000155; the same query repeated hit the cache in 0 ms. This is an integration test, not a quality benchmark.

Laya in the Jev project

Local alternative under evaluation

Laya is a candidate for comparing local structured decisions with Jev. References and analysis are available; the Jev adapter and bot v3 integration remain pending.

What we have checked

In the local analysis on September 21, 2026, 14 Laya application tests passed and 16 previously saved responses were accepted by the Jev structural validator. No new inference was performed in that analysis.

The existing educational report records 13 correct answers out of 16 synthetic Portuguese examples, including high-confidence errors. This does not establish superiority over Jev or production quality.

Next step: a comparable pilot

Use the same cases and criteria, separate tuning from testing, measure quality, full cost and end-to-end latency, and retain human review. Laya confidence has its own meaning; do not automatically reuse Jev thresholds.

What has already been observed — and what still needs measuring

Simulation for learning

The 20 public cases and the instructional replay are original and simulated. They are meant to study phrasing, policies and errors; they do not measure model quality.

View the instructional comparison

Rule baseline

The report records 20 correct out of 24 fictitious tickets (83.33%), with a macro‑F1 of 0.84235. These results come from the rules, not from Jev.

Examine the rule report

Real integration tested

On 09/19/2026, the ten packages received responses via OpenRouter, without failures and with the expected classifications on the fictitious examples. This confirms the integration on those examples; it does not constitute an independent benchmark.

Read the report of the ten queries

Evaluation still needed

Quality in Portuguese, calibration and operational use require independent data and human reference. The project does not perform payments, diagnostics, merges or agent actions.

Read exaggerations, doubts and limits

The seven new packages were verified with fixtures and controlled tests. The real inference report covers only the original ten packages.

Complete project and course library

Public documentation to dive deeper into each part. The repositories and study materials are in Portuguese; this page organizes the access paths.

Choose your next step.

Try a case, study the corresponding module, or adapt a package to your project.