← Automation Arsenal
Field notes Stage 02 · Execute
Case study · agentic browser testing

Browser Usegoals, not selectors

Browser Use is an open-source Python library that lets an LLM agent drive a real browser. Instead of scripting every click, you give it a goal. I use it for the checks scripts are bad at: exploratory runs, long-tail flows, and "can a real user actually do this?"

What it is
Open-source library for LLM browser agents
Language
Python, async
Finds elements by
Meaning on the page, not CSS selectors
Best for
Exploratory and goal-based checks
How it works

An observe → think → act loop

The agent repeats one cycle until the goal is met, it gives up, or it hits the step limit.

1

Observe

The page's interactive elements are extracted and numbered, so the model can refer to them.

2

Think

The LLM compares the page with the goal and picks the next action, such as "click element 4".

3

Act

The action runs in a real browser: click, type, scroll, navigate or open a tab.

4

Verify

My addition: a deterministic check confirms the outcome, because an agent's "done" is a claim, not proof.

Interactive · agent loop

Watch the agent shop, then ship a redesign

Run it once. Then flip "Ship a UI redesign" and run again: the selector script breaks, the agent adapts, and the final check still decides.

agent run · stagingSimulated agent · mock storefront
Task Add the cheapest USB-C cable to the cart and report the cart total. allowed: staging.shop.localmax_steps: 25test account only
https://staging.shop.local/
Agent logstep 0 / 25

    The code

    A goal, a model, and a check that doesn't trust the agent

    The agent does the exploring. A plain assertion against the system of record decides pass or fail.

    
    import asyncio
    from browser_use import Agent, ChatAnthropic  # the API moves fast; check the current README
    
    TASK = (
        "Open https://staging.shop.local, add the cheapest USB-C cable "
        "to the cart, and report the cart total."
    )
    
    async def main():
        agent = Agent(task=TASK, llm=ChatAnthropic(model="claude-sonnet-5"))
        # Guardrails: staging only, a throwaway test account, and a hard step cap
        history = await agent.run(max_steps=25)
        print(history.final_result())
    
    asyncio.run(main())
    
    Seven lines to a working agent. The guardrails matter more than the code.
    Where it fits

    Stage 02 of my quality pipeline

    It runs alongside the scripted suites. It doesn't replace them.

    Trade-offs

    Powerful, and it needs guardrails

    Strengths

    • Survives UI churn. It targets intent, so a redesign doesn't break it the way selectors break.
    • Finds the long tail. It's good for exploratory passes over flows nobody wrote a script for.
    • Works on legacy apps with no test IDs and no page objects.
    • Tasks read like requirements, so non-engineers can follow along.

    Watch-outs

    • Non-deterministic. The same task can take a different path, so it's not a per-PR regression gate.
    • Slower and costs model tokens on every step, compared with a script.
    • Needs guardrails: allowed domains, test accounts only, step caps, no production data.
    • "Done" is a claim. Always confirm the outcome with a deterministic check.
    My rule of thumb

    Agents explore. Scripts assert.

    Use the agent to discover and adapt, and plain code to decide. That keeps the flexibility without giving up trust.

    • Agentic runs are nightly or exploratory, not the PR gate.
    • Every run is scoped to staging with a test account.
    • Outcomes are checked against the backend, not the agent's summary.
    • Anything the agent finds twice becomes a scripted test.