Overview
Evaluate a model's ability to implement AI Application QA Agent as a bounded, evidence-backed product challenge.
Capabilities tested
- browser automation
- QA
- evidence
Required outcome
A URL-driven QA tool that reports interaction, console, request, mobile, keyboard, and accessibility failures with evidence.
Requirements
functional
- A URL-driven QA tool that reports interaction, console, request, mobile, keyboard, and accessibility failures with evidence.
ux
- Responsive, readable interface
- Keyboard-accessible core controls
technical
- Production build succeeds
- No secrets or private data in the artifact
Constraints
Allowed
- Local sample data
- Established framework utilities
Forbidden
- Automated model execution
- Unreviewed private data
- Unsupported production claims
Acceptance criteria
- The primary workflow is complete and observable
- Core controls change real local state
- Failure or empty states are visible
- The narrow viewport remains usable
Evidence required
- Production build output
- Desktop and mobile screenshots
- Core-flow recording
- Known-bug list
Exact prompt
prompts/challenges/ai-app-qa-agent.md
# BuildArena Challenge Role: Senior product engineer. Project: AI Application QA Agent v1.0.0 Objective: Evaluate a model's ability to implement AI Application QA Agent as a bounded, evidence-backed product challenge. Required outcome: A URL-driven QA tool that reports interaction, console, request, mobile, keyboard, and accessibility failures with evidence. Functional requirements: - A URL-driven QA tool that reports interaction, console, request, mobile, keyboard, and accessibility failures with evidence. UX requirements: - Responsive, readable interface - Keyboard-accessible core controls Technical requirements: - Production build succeeds - No secrets or private data in the artifact Constraints: - Allowed: Local sample data; Established framework utilities - Forbidden: Automated model execution; Unreviewed private data; Unsupported production claims Build order: Inspect the environment; establish the smallest vertical slice; verify the central workflow; add secondary requirements; run production validation. Evidence to return: Commands, test results, screenshots, core-flow recording, known bugs, missing requirements, human intervention, final commit, and deployment URL only when authorized. Claims prohibited without proof: complete, production-ready, fully working, bug-free, accessible, responsive, performant.