
The Hostile Web: Securing Browser Agents
A browser agent operates inside pages that may be misleading, compromised, or designed to redirect its behavior. The web must remain data—not authority.
Read MoreZharfAI Team

Computer-use agents look impressive when a task has one clean path. Real interfaces contain stale sessions, pop-ups, ambiguous labels, partial saves, permissions, and data that moves while the agent is working.
Evaluation should reproduce that reality.
Define the desired end state and unacceptable side effects. For an expense workflow, success may mean the correct receipt is attached, fields match the source, policy exceptions are flagged, and nothing is submitted without approval. A sequence of plausible clicks is not enough.
Build cases across four dimensions:
Record the observation available at each step, the action selected, tool response, and resulting state. Mask sensitive data, but keep enough structure to replay a failure. Without step-level evidence, teams cannot distinguish perception errors from bad planning or unsafe execution.
An agent that recognizes uncertainty and asks for help may be more valuable than one that completes more tasks by guessing. Track safe abandonment, successful resume, duplicated actions, and quality of escalation alongside completion rate.
The best test suite grows from production incidents and difficult human examples. A task bench is not a one-time leaderboard; it is the regression system that keeps an evolving agent dependable.

A browser agent operates inside pages that may be misleading, compromised, or designed to redirect its behavior. The web must remain data—not authority.
Read More
Human-in-the-loop design works when approval is reserved for consequential uncertainty and presented with enough evidence to make a real decision.
Read More
Before launch, an AI feature needs an owner, evaluation gates, security boundaries, observability, cost limits, fallback, and controlled change.
Read MoreGet in touch with our team to discuss how we can help your business.