AI-driven browser automation that understands intent, not selectors. No XPath. No CSS queries. Just natural language — and a Nova LLM reasoning through every step like a senior QA engineer.
Nova Act is Amazon's agent SDK that converts plain-English instructions into real browser actions — powered by a multimodal LLM that sees, reasons, and acts.
Traditional automation (Selenium, Playwright scripts, Cypress) requires engineers to write brittle locators that break on every UI change. Maintaining them costs more than writing them. Nova Act eliminates this entirely.
Instead, you describe what you want in plain English: "Open amazon.com, search for laptop, filter by 4 stars." The Nova LLM interprets the intent, decomposes it into atomic steps, and executes each one via Playwright against a real Chromium browser.
After every action, the agent captures a screenshot + DOM snapshot, reasons about the result, and decides the next move. It's not a script. It's an autonomous agent that thinks.
from nova_act import NovaAct # That's it. One line. The agent handles everything else. with NovaAct(starting_page="https://amazon.com") as nova: nova.act("Search for laptop and click the first result")
From browser input to AI reasoning to Playwright execution — the complete request lifecycle of a Nova Act agent session.
# Conceptual model of Nova Act's internal observe-reason-act loop def agent_loop(instruction: str, browser) -> Result: goal_achieved = False actions_taken = [] while not goal_achieved: # ① OBSERVE — capture current page state screenshot = browser.screenshot() dom_tree = browser.accessibility_tree() # ② REASON — LLM decides next action decision = nova_llm.reason( instruction=instruction, screenshot=screenshot, dom=dom_tree, history=actions_taken ) if decision.type == "done": break # ③ ACT — Playwright executes the command browser.execute(decision.action) actions_taken.append(decision.action) return Result(actions=actions_taken, status="success")
Type any automation task in natural language. Watch the Nova Act agent reason through it — step by step, action by action. Connect your FastAPI backend for live browser execution.
A FastAPI backend orchestrating Nova Act sessions, served through a minimal Jinja2 frontend — connected to AWS Bedrock via IAM credentials.
fastapi==0.110.0 uvicorn==0.29.0 nova-act>=0.1.0 playwright>=1.44.0 jinja2==3.1.4 boto3>=1.34.0 # S3 logging pydantic==2.7.1
from fastapi import FastAPI, Request from fastapi.templating import Jinja2Templates from fastapi.responses import JSONResponse from pydantic import BaseModel from nova_act import NovaAct import uvicorn app = FastAPI(title="Nova Act Web Agent") templates = Jinja2Templates(directory="templates") class AgentRequest(BaseModel): instruction: str starting_page: str = "https://google.com" @app.get("/") async def index(request: Request): return templates.TemplateResponse( "index.html", {"request": request} ) @app.post("/run-agent") async def run_agent(body: AgentRequest): try: with NovaAct( starting_page=body.starting_page, logs_directory="./logs" ) as nova: result = nova.act(body.instruction) return JSONResponse({ "status": "success", "actions": result.actions_taken, "url": result.final_url, }) except Exception as e: return JSONResponse( {"status": "error", "message": str(e)}, status_code=500 ) if __name__ == "__main__": uvicorn.run(app, host="0.0.0.0", port=8000)
Six fundamental advantages of replacing selector-based scripts with an AI agent that thinks, adapts, and heals.
From autonomous CI/CD gates to multi-agent orchestration — the next evolution of AI-driven QA is already taking shape. These are the extensions I'm building toward.