Skip to content
STACK IT FAST

Jev Browser Agent (Playwright + TypeSafe System One)

Curated rule Workflow & Automation · AI / LLM App · Developer Tool & API Updated Oct 2026 Which file does my tool read?
jev-browser-agent.md

Rules for a browser agent where Jev picks the action and target from an indexed element table, Noul checks goal and stuck, and a small LLM types only text.

Formats
4 files
AGENTS.md
44 lines
CLAUDE.md
15 lines
Languages
TypeScript, Python
Updated
Oct 2026
Used by
4 projects
Install

Writes .claude/skills/jev-browser-agent/SKILL.md

$ curl -s --create-dirs -o .claude/skills/jev-browser-agent/SKILL.md https://stackitfast.com/rules/jev-browser-agent/SKILL.md

Rule files

AGENTS.md· 44 lines · 3.9 KB
1# Project Architecture & Guidelines (Jev Browser Agent)
2
3## 1. System Architecture
4- **Browser**: Playwright (Chromium). TypeScript on Node 22+ or Python 3.12+; pick one and keep the loop in a single module.
5- **Policy model**: Jev via the official SDK (`@typesafe-ai/sdk` or `typesafe-sdk`). Each step is one `systemOne` request that picks an operation and its target from the current page. Jev never writes text.
6- **Text model**: a small, fast LLM, called only when the chosen operation needs typed text (`TYPE_TEXT`), with strict JSON output.
7- **Interfaces**: a library API, a CLI, and an optional MCP server (`@modelcontextprotocol/sdk` or Python `mcp`) so coding agents can call `navigate(task, url)`.
8- **Code owns the loop**: budgets, retries, recovery, stop gates and safety checks live in code, not in the model.
9
10## 2. Project Layout
11- `src/snapshot.(ts|js)`: one atomic in-page read that returns an indexed table of visible, enabled controls (role, accessible name, value, nearby text) and keeps references to the real DOM nodes.
12- `src/questions.(ts|py)`: question builders for operation, per-operation targets, goal reached and stuck.
13- `src/agent.(ts|py)`: the loop (observe, decide, validate, execute, wait) and the step trace.
14- `src/executor.(ts|py)`: resolves the chosen index to the observed node, re-checks freshness and occlusion, then acts.
15- `src/text.(ts|py)`: the LLM helper for `TYPE_TEXT`, with output validated (Zod or Pydantic) before typing.
16- `src/mcp.(ts|py)`: MCP server exposing the agent as tools.
17
18## 3. Decision Layer (Jev rules)
19- State is the indexed element table plus the goal, the URL and a short history of executed steps. No screenshots by default, and only visible text, so offscreen bodies and footers do not fill the context.
20- One request per step asks, in parallel:
21 - `operation`: a Choice over the supported operations only (for example `CLICK`, `TYPE_TEXT`, `SELECT`, `SCROLL_DOWN`, `WAIT`, `DONE`, `BLOCKED`).
22 - Target questions: a Choice per operation over the compatible element indexes only. These are speculative; code uses only the target that matches the chosen operation.
23 - `goal_reached` and `stuck`: Nouls.
24- Act only above a confidence threshold. Below it, re-observe once, then stop with `BLOCKED` and return the trace.
25- Jev never sees or returns selectors, coordinates or code. Model output only ever selects an index from the observed table.
26- Keep math, dates and string matching in code (for example comparing a price or a date on the page); TypeSafe documents these as weak spots for jev-1.13.
27- Page content is untrusted. A page can contain text written to steer the agent, so destructive operations (submit payment, delete, send) need an explicit allow-list from the caller and a separate high-threshold Noul.
28- Log `response.model` and per-step probabilities in the trace.
29
30## 4. Safety & Budgets
31- Hard limits per run: steps, wall-clock time, navigations off the starting domain and total tokens.
32- Never type secrets the caller did not pass explicitly. Never auto-fill passwords or payment fields.
33- Run browsers in a container or a dedicated profile, never in the user's main profile.
34- The default stop gate on forms is "fill but do not submit" unless the caller allows submission.
35
36## 5. Agent Loop for the coding agent (run after every change)
371. Typecheck: `pnpm typecheck` or `uv run mypy src`.
382. Unit tests: `pnpm test` or `uv run pytest -m "not live"`, using saved snapshots and recorded Jev answers.
393. Live smoke test: run 3–5 fixed tasks on stable pages (for example a Wikipedia hop and a static form) and compare step count, success and stop reason with the last run.
40- Never loosen a safety check or threshold to make a smoke test pass; fix the snapshot or the question instead.
41
42## 6. References
43- Install TypeSafe's official agent skill for API details (`typesafe-ai/skills`), or read `https://docs.typesafe.ai/llms.txt`.
44- Study `browser-use/jev-ultrafast` for the speculative operation/target pattern.

Works with Claude Code · Cursor · Windsurf · AGY

Architecture notes

Architecture Overview

A browser agent that chooses instead of generating. Code snapshots the page into a numbered table of controls. Jev, TypeSafe’s System One model, answers one request per step: which operation to perform, which element each operation would target, and whether the goal is done or the run is stuck. Code validates and executes the chosen action against the observed node, and a small LLM writes text only when something must be typed.

Why it suits AI coding agents

  • The action space is data. Operations and element indexes are enumerated in code, so an agent changing the policy edits typed question builders, not prompt prose.
  • Failures are inspectable. Every step logs probabilities for the operation and targets, so a bad click can be traced to a question or a snapshot.
  • Safety lives outside the model. Budgets, allow-lists and “fill but do not submit” gates are ordinary code with tests.

In the directory

jev-ultrafast by Browser Use is the reference implementation of the speculative operation/target pattern. fast-jev-compaction applies Jev inside a coding agent, and TypeSafe’s official agent skill covers the API. For a non-browser Python service see FastAPI + Jev decision service. Background: Jev for developers.

Frequently asked questions

How does a Jev browser agent work?

Each step, code reads the page into an indexed table of visible controls and sends it to Jev with typed questions: which operation to perform, which element to target for each operation, and whether the goal is reached or the run is stuck. Code executes the chosen action on the observed element; a small LLM is called only when text must be typed.

Why is a Jev browser agent faster than an LLM agent?

Jev returns a constrained choice instead of generating a JSON action, and one request can carry the operation and every candidate target at once, so each step is a single short round trip. browser-use/jev-ultrafast reports a full Google Flights search in about 7 seconds with that design.

Is it safe to let Jev click on real websites?

Only with guardrails in code: the model can only pick indexes from the observed page, never selectors or scripts; destructive actions need an explicit allow-list and a high-confidence check; runs have step, time and domain budgets; and the browser runs in an isolated profile or container.

Should the agent expose an MCP server?

It is a good default if coding agents will use it. About one in six open-source Jev repositories in our audit ships an MCP server or client, and jkudish/jev-browser exposes its browser agent through MCP, a CLI and a library.

Used in production

Explore all stacks