Amazon Nova Act◆ AI Browser Automation◆ Aayush Mishra◆ SDET · AI-Driven Testing◆ Natural Language → Browser Control◆ FastAPI + Playwright + AWS Bedrock◆ Self-Healing Automation◆ Amazon Nova Act◆ AI Browser Automation◆ Aayush Mishra◆ SDET · AI-Driven Testing◆ Natural Language → Browser Control◆ FastAPI + Playwright + AWS Bedrock◆ Self-Healing Automation◆
Amazon Nova Act · 2024

BROW­SER AGENTS THAT THINK

AI-driven browser automation that understands intent, not selectors. No XPath. No CSS queries. Just natural language — and a Nova LLM reasoning through every step like a senior QA engineer.


01 What is Nova Act

The AI that drives your browser.

Nova Act is Amazon's agent SDK that converts plain-English instructions into real browser actions — powered by a multimodal LLM that sees, reasons, and acts.

From instruction to action — without writing a single selector.

Traditional automation (Selenium, Playwright scripts, Cypress) requires engineers to write brittle locators that break on every UI change. Maintaining them costs more than writing them. Nova Act eliminates this entirely.

Instead, you describe what you want in plain English: "Open amazon.com, search for laptop, filter by 4 stars." The Nova LLM interprets the intent, decomposes it into atomic steps, and executes each one via Playwright against a real Chromium browser.

After every action, the agent captures a screenshot + DOM snapshot, reasons about the result, and decides the next move. It's not a script. It's an autonomous agent that thinks.

nova_act_hello.py
from nova_act import NovaAct

# That's it. One line. The agent handles everything else.
with NovaAct(starting_page="https://amazon.com") as nova:
    nova.act("Search for laptop and click the first result")
🧠
LLM Reasoning
Nova LLM receives screenshot + DOM and reasons about which element to interact with — no selectors needed at any step.
🔄
Observation Loop
After every action, the agent observes the new page state before planning its next move — adapting dynamically to any result.
🎭
Playwright Core
Real Chromium automation under the hood. Full JS execution, network access, multi-tab support.
02 Paradigm Shift

THE
GREAT SHIFT
IN TESTING.

Legacy
Selenium / Playwright Scripts / Cypress
  • ✕ CSS selectors & XPath required for every element
  • ✕ Breaks on every UI refactor or class rename
  • ✕ Manual script authoring by senior engineers only
  • ✕ High maintenance cost — tests rot over time
  • ✕ Zero contextual or semantic understanding
  • ✕ Business users cannot write or read test cases
  • ✕ Flaky timing issues and brittle explicit waits
→
AI-Native
Amazon Nova Act · Agent SDK
  • ✓ Plain English instructions — no selectors ever
  • ✓ Self-heals automatically when UI changes
  • ✓ Anyone can write tests — no code required
  • ✓ Near-zero maintenance — AI adapts each run
  • ✓ Semantic + visual page understanding
  • ✓ PMs and QAs can own test scenarios
  • ✓ Intelligent waits — agent knows when to proceed
03 System Architecture

Nine layers,
one thought.

From browser input to AI reasoning to Playwright execution — the complete request lifecycle of a Nova Act agent session.

1
User Interface
HTML/CSS/JS — instruction input, real-time log display
2
FastAPI Backend
Python server — POST endpoint, validation, async runner
3
Nova Act SDK
Core Python SDK — session management, agent lifecycle
4
Nova LLM (Bedrock)
Amazon's vision-language model — intent interpretation
5
Agent Planner
Decomposes goal → ordered atomic action sequence
6
Browser Controller
Routes commands to Playwright executor
7
Playwright Engine
Real Chromium — click, fill, navigate, scroll, wait
8
Target Web App
The live site being automated
9
Logs + Results
Streamed to UI · archived to S3
Layer 1–2 · Frontend + API
The Entry Gate
The browser UI accepts plain-English instructions and sends them via HTTP POST to FastAPI. The server validates the payload, extracts the instruction and starting URL, then spawns an async Nova Act session — streaming logs back to the client in near real-time.
Layer 3–4 · SDK + LLM
The Intelligence Core
The Nova Act SDK initializes a browser session and sends the instruction to Amazon's Nova LLM via AWS Bedrock. The multimodal model receives the instruction alongside a screenshot and DOM snapshot — reasoning about intent like a senior QA engineer reading a test brief.
Layer 5–7 · Planner + Browser
The Execution Engine
The agent planner converts LLM output into a typed action sequence: navigate → click → fill → press Enter. Each action maps to a Playwright API call against a real Chromium instance. After each step, a new screenshot is captured and fed back to the LLM — closing the observe-reason-act loop.
Layer 8–9 · Target + Output
Results & Observability
Execution logs, screenshots, action traces, and final page URL are returned to the frontend and optionally archived to Amazon S3. Each session produces a full audit trail — perfect for CI/CD gates, QA reports, or debugging complex agent behaviour.
04 Agent Internals

THE AGENT
THINKS.

01
📝
Instruction
Plain English task received and tokenized by Nova LLM
02
🧠
LLM Reasons
Screenshot + DOM analyzed — intent resolved, next action chosen
03
📋
Plan Action
Atomic typed command generated — click, fill, navigate, scroll
04
🖱️
Execute
Playwright fires the command in real Chromium browser
05
👁️
Observe ↺
New screenshot captured → loop repeats until goal complete
agent_loop_concept.py
# Conceptual model of Nova Act's internal observe-reason-act loop

def agent_loop(instruction: str, browser) -> Result:
    goal_achieved = False
    actions_taken = []

    while not goal_achieved:
        # ① OBSERVE — capture current page state
        screenshot = browser.screenshot()
        dom_tree   = browser.accessibility_tree()

        # ② REASON — LLM decides next action
        decision = nova_llm.reason(
            instruction=instruction,
            screenshot=screenshot,
            dom=dom_tree,
            history=actions_taken
        )

        if decision.type == "done":
            break

        # ③ ACT — Playwright executes the command
        browser.execute(decision.action)
        actions_taken.append(decision.action)

    return Result(actions=actions_taken, status="success")
05 Interactive Demo

AGENT
PLAYGROUND.

Type any automation task in natural language. Watch the Nova Act agent reason through it — step by step, action by action. Connect your FastAPI backend for live browser execution.

Agent Instruction
Quick Presets
nova-act · agent session log
idle
00:00:00INFNova Act agent initialized · Awaiting instruction · Press ▶ to begin
⚠ REGION NOTE — The Nova Act visual Playground may be unavailable in certain regions (India included). Use the Nova Act Python SDK directly via CLI for full functionality. This demo simulates the agent reasoning flow end-to-end.
06 My Implementation

HOW I BUILT
THIS PROJECT.

A FastAPI backend orchestrating Nova Act sessions, served through a minimal Jinja2 frontend — connected to AWS Bedrock via IAM credentials.

Project Structure
nova-act-web-agent/
│
├── app.py← FastAPI + Nova Act runner
├── templates/
│ └── index.html← Jinja2 frontend
├── static/
│ └── style.css← Stylesheet
└── requirements.txt← Python deps
Requirements
requirements.txt
fastapi==0.110.0
uvicorn==0.29.0
nova-act>=0.1.0
playwright>=1.44.0
jinja2==3.1.4
boto3>=1.34.0    # S3 logging
pydantic==2.7.1
Backend — app.py
app.py
from fastapi import FastAPI, Request
from fastapi.templating import Jinja2Templates
from fastapi.responses import JSONResponse
from pydantic import BaseModel
from nova_act import NovaAct
import uvicorn

app = FastAPI(title="Nova Act Web Agent")
templates = Jinja2Templates(directory="templates")

class AgentRequest(BaseModel):
    instruction: str
    starting_page: str = "https://google.com"

@app.get("/")
async def index(request: Request):
    return templates.TemplateResponse(
        "index.html", {"request": request}
    )

@app.post("/run-agent")
async def run_agent(body: AgentRequest):
    try:
        with NovaAct(
            starting_page=body.starting_page,
            logs_directory="./logs"
        ) as nova:
            result = nova.act(body.instruction)
        return JSONResponse({
            "status":  "success",
            "actions": result.actions_taken,
            "url":     result.final_url,
        })
    except Exception as e:
        return JSONResponse(
            {"status": "error", "message": str(e)},
            status_code=500
        )

if __name__ == "__main__":
    uvicorn.run(app, host="0.0.0.0", port=8000)
07 Setup Guide

UP AND
RUNNING IN
MINUTES.

01
Create Virtual Environment
Isolate dependencies in a clean Python environment before installing anything.
python -m venv venv && source venv/bin/activate
02
Install Dependencies
Install FastAPI, Nova Act SDK, Playwright, and all Python packages.
pip install -r requirements.txt
03
Install Playwright
Download the Chromium browser binary that Nova Act will control.
playwright install chromium
04
Create IAM User
AWS Console → IAM → Create User → Attach AmazonBedrockFullAccess → generate Access Key.
AWS Console · IAM · Bedrock Access
05
Configure AWS Credentials
Set your AWS access key and secret — Nova Act uses Bedrock via these credentials.
aws configure
06
Start FastAPI Server
Launch the backend in development mode with hot reload enabled.
uvicorn app:app --reload --port 8000
07
Open in Browser
Navigate to the local server, type an instruction, click Execute Agent.
http://localhost:8000
08
Run Your First Agent
Type any instruction in plain English and watch the Nova Act agent execute it live.
"Search for laptop on amazon.com"
08 Why AI Automation

THE CASE
FOR AGENTS.

Six fundamental advantages of replacing selector-based scripts with an AI agent that thinks, adapts, and heals.

01
⚡
Zero Selector Maintenance
No more updating XPath and CSS selectors every sprint. The AI finds elements by semantic meaning and visual appearance — surviving any UI refactor automatically.
Core Advantage
02
✍️
Plain English Test Scenarios
Business analysts, product managers, and QAs can author complete test cases in plain English. No Gherkin syntax, no code review, no engineering bottleneck.
Accessibility
03
🔧
Self-Healing Behaviour
When a page changes, the agent adapts its interaction strategy in real-time instead of failing with NoSuchElementException. Fewer broken pipelines. No more 2 AM pager alerts.
Resilience
04
🚀
10× Faster Test Authoring
What takes hours of selector debugging reduces to a single sentence. Dramatically compresses time-to-coverage for new features in fast-moving sprint environments.
Velocity
05
🤖
AI-Assisted QA Pipelines
Pair Nova Act with an LLM that generates test cases from requirements docs. Close the loop from spec → test generation → execution → report — entirely autonomously.
AI Workflow
06
🌐
Democratised Quality
SDET teams are no longer bottlenecks. Any team member who can describe a user flow in English can create and run automated tests — making quality a shared responsibility.
Team Scale
09 Future Extensions

WHAT'S NEXT.

Where this project goes from here.

From autonomous CI/CD gates to multi-agent orchestration — the next evolution of AI-driven QA is already taking shape. These are the extensions I'm building toward.

Built with
Python FastAPI Nova Act SDK AWS Bedrock Playwright Amazon S3
🔧
Self-Healing Test Suites
When a test fails, an LLM analyzes the failure, proposes a fix, and applies it automatically — creating a continuously self-maintaining test corpus with minimal human intervention.
✨
AI-Generated Test Cases
Feed product requirements to an LLM that auto-generates comprehensive test scenarios, edge cases, and acceptance criteria — then executes them immediately via Nova Act.
👁️
Visual Regression Testing
Combine Nova Act screenshots with a vision model to detect unintended visual changes across releases — pixel-perfect diffing powered by AI semantic understanding.
🚀
CI/CD Automation Gates
Embed Nova Act agents in GitHub Actions or Jenkins pipelines — autonomous QA gates that run, evaluate, and block merges before every production deploy.
🕸️
Multi-Agent Orchestration
Deploy parallel specialized agents coordinated by a supervisor for full E2E coverage at scale — checkout agent, search agent, auth agent all running simultaneously.
📱
Mobile + API Coverage
Extend the architecture to Appium for mobile browsers and REST API contract testing — one unified AI agent layer across web, mobile, and API surfaces.