Readiness Check methodology

We test actions, not impressions.

Each task runs three times against the live website. A separate AI judge reviews the evidence before any result becomes a finding.

Public market signalSeptember 2026

The customer may now arrive through an agent.

Meta’s September 8 launch of Muse provides a current example: a personal AI agent with its own browser that can navigate websites, fill forms, book appointments, and—with the user’s approval—make purchases. That makes task completion by agents a live website concern rather than a theoretical one.

ViaLayer tests the broader condition this creates: whether an outside agent can understand a business and complete a defined customer path on its public website.

Muse is context, not a ViaLayer integration or test target. Results describe the systems and tasks named in each report.

One task. Three runs. Independent review.

The check uses the same public paths available to a customer. No private access, prepared demo, or cooperation from the site is used during a run.

01

Define practical customer tasks.

Tasks match the business and use plain language. Typical examples include finding current hours, confirming a service area, requesting a quote, or booking an appointment.

02

Run against the live site.

An AI agent receives a real browser session and attempts each task through the public website. The run records steps, site responses, and screenshots.

03

Repeat in separate sessions.

Every task starts again twice. Three independent runs distinguish a repeatable site problem from ordinary variation in one agent run.

04

Review the complete evidence.

A separate AI judge reviews the transcript and evidence for every run. The result does not rely on the acting agent's self-reported success.

05

Classify without smoothing the result.

The report states the outcome, consistency, supporting evidence, and confidence classification. Mixed results remain mixed.

Confidence is shown, not implied.

Every finding carries an explicit class and the exact result across the three independent runs.

Hard finding

Repeatable and directly supported.

The same material outcome occurs across the runs and is supported by a site response, transcript, or screenshot. The exact count appears beside it.

Soft finding

Useful, but less clear-cut.

The result varies across runs or depends on interpretation. The report states the mixed outcomes and why confidence is limited.

Inconclusive

Not enough evidence.

The available runs do not support a reliable conclusion. We explain what prevented classification and do not present a confirmed problem.

What the report records.

Every outcome can be traced to an attempted task and its evidence.

OutcomePass, fail, timeout, or inconclusiveThe observed result without concealing mixed runs
ConsistencyExact count across three runsFor example, 3 of 3 fail or 2 fail and 1 timeout
EvidenceTranscript, site response, and screenshotsThe material reviewed to support the result
Independent reviewSeparate AI judge decisionA second assessment apart from the acting agent

What the check does not claim.

It is not a prediction of every AI system.

A Readiness Check tests defined tasks at a specific point in time. Different systems and customer prompts can behave differently.

One run is not treated as certainty.

Agent behavior varies. The check repeats each task and reports mixed outcomes rather than converting them into a cleaner story.

A past result is not permanent assurance.

Websites, directories, forms, and agent behavior change. Monitoring is required when ongoing stability matters.

Responsible testing is part of the method.

Public checks use public paths.

No administrative access is needed for the initial Readiness Check. Any later access is separately scoped, limited, and revocable.

The stopping point is defined first.

Testing stops before payment, contract acceptance, account changes, destructive actions, or other irreversible steps unless a separate written procedure explicitly authorizes them.

Controlled submissions require authorization.

If a meaningful test needs a form or booking submission, the identifier, allowed details, expected handling, and cleanup responsibility are agreed in advance.

Unexpected sensitive steps stop the run.

A payment, legal, authentication, sensitive-data, or security boundary is reported rather than crossed.

Read the complete Terms and Responsible Testing policy →

See the method in a report.

Review a wholly fictional example, or ask ViaLayer to test one practical task on your live website.