jev agent turns browser automation into exactly that kind of question: at every step it shows Jev the controls on the current page and asks which one to act on next. A regular chat model is called only when a value has to be typed into a field.
In Anchor Browser, this agent is available through the agentic APIs with taskOptions.agent = 'jev'.
Overview
The Jev agent provides:- Fast steps — one classification call per action, no screenshots and no free-form reasoning
- Choices, not guesses — Jev can only pick from the controls that are actually on the page, so every action lands on a real element
- Text generation only when needed — typed values come from your configured
provider/model(or BYOM); every other step is pure classification - Lean token usage — classification returns a choice instead of text, and the text model runs just for the fields you fill
Code Example
How Jev Drives the Task
Jev is a classifier, not a planner. It never writes selectors, coordinates, or code. The agent owns the loop and uses Jev for one thing: deciding the next move from a closed list of options.1
Observe the page
The agent reads the current page and builds a numbered list of the controls a user could act on right now: links, buttons, text fields, checkboxes, dropdown options, and so on, each with its label and current state. It also captures the visible text so Jev can tell whether the goal is already satisfied.
2
Ask Jev for the next move
Jev receives the goal, the page summary, the control list, and the recent action history, and returns a single decision:
CLICK,TYPE_TEXT, orSELECTon one specific control from the listSCROLL_DOWN,SCROLL_UP, orWAITwhen the needed control is not available yetDONEwhen every requirement in the goal is visibly satisfiedBLOCKEDwhen no offered action can make progress
3
Generate the value for typed fields
Only
TYPE_TEXT needs text. The configured text model receives the goal, the selected field’s label and current value, and the page context, and returns the exact string to type. Secret placeholders such as <secret>EMAIL</secret> are kept verbatim and swapped for the real value at type time. If no value can be determined, the task stops as blocked instead of typing a guess.4
Act and repeat
The agent performs the chosen action on the real element, confirms the page has not changed since Jev decided (and re-asks if it has), then observes again. The loop continues until Jev reports
DONE or BLOCKED, or the step limit is reached.Example walkthrough
The Wikipedia task above completes in three actions and one closing decision:
Each row is one decision. The text model ran once, at step 1.
Compared with the default agent
The same task, run once on each agent withmax_steps: 30:
Results vary with the site and the task. Jev’s advantage is largest on click-through and form flows, where most steps need no generated text.
Configuration Options
Jev is in beta and currently focuses on navigating and operating pages. Structured output (
output_schema), human-in-the-loop (human_intervention), OS-level control (use_os_control), and element detection (detect_elements) are not available yet, so leave them unset when choosing this agent. use_action_index_tools is accepted and has no effect.Which model does what
With BYOM active, all BYOM providers are supported for typed text (
openai, anthropic, google, vertex, azure, custom) and model is required.
Result
The task returns plain text: a status line followed by the visible text of the final page.Stopped (blocked) at …. Reported token usage includes both the classification calls and the text model.
When the Agent Stops
- Jev answers
DONE— every requirement is visibly satisfied on the current page - Jev answers
BLOCKED— no offered action can make progress max_stepsis reached- Several consecutive actions did not change the page
- The text model could not determine a value for a required field
- The page kept changing before an action could be confirmed
Limitations
- Decides, does not write — Jev returns a choice, never text. The agent operates the page; it does not summarize, extract, or answer questions, and the result is the final page state. Use another agent when the deliverable is generated content.
- Needs real controls — Jev picks from the labeled controls a page exposes. Purely visual interfaces such as canvas apps, maps, games, and image CAPTCHAs offer nothing to pick from, so use a screenshot-based agent there.
- One step at a time — every decision is made from the current page and recent history. Tasks that hinge on comparing content across pages or interpreting ambiguous instructions are better served by a reasoning agent.
- Done means visible — Jev reports
DONEonly from what is on screen. Outcomes the page does not visibly confirm cannot be verified by the agent.
Secure Credentials with Secret Values
Secret values are substituted at type time and never sent to Jev or the text model.Best Practices
- State the visible end condition in the goal (for example, “stop when the order confirmation is visible”). Jev only answers
DONEwhen the page shows that every requirement is met. - Prefer Jev for form-driven sites such as search, filters, and multi-page flows with native controls. Use a computer-use agent for canvas or purely visual interfaces.
- Pick a fast text model — it is called only for
TYPE_TEXT, so a small model keeps steps quick without affecting decision quality.

