Design Repeatable Tests in Prompt Lab
Use LLMIC Prompt Lab to design repeatable AI tests with fixed questions, provider settings, evidence and review criteria.
Good prompt research changes one thing at a time and records enough context to repeat the observation later.
Build availability: This guide describes the September 2026 development workflow. Some controls may not yet be included in the public installer. Check the release notes and your app version before following the steps.
Evidence flow
- State the questionWrite the decision the test should inform.
- Create a fixed prompt setVersion wording, audience and language.
- Control the runKeep provider, model, search and location settings.
- Review the evidenceInspect answers, citations, failures and variation.
This visual explains the review sequence. It is not a performance forecast or a set of measured customer results.
Start with a clear question
Open the relevant LLMIC dashboard only after loading the crawl, property or saved observation needed for this review. For AI prompt testing, record the website, date, filters and document layer before interpreting a result. This keeps the work reproducible when another team member reviews it later.
Read the evidence before the recommendation
Open the affected page or row and inspect the retained evidence. Counts help you find patterns, but the page-level source explains what was actually measured. When coverage is partial, a response failed, or a legacy crawl omitted the required field, keep that state visible. Do not turn unavailable evidence into a failed check.
Choose the smallest useful action
Write an instruction that names the exact URL, observed condition and intended result. Preserve intentional behavior and review shared patterns before applying a site-wide change. If a suggestion creates new wording, markup or redirects, a person should approve the facts and destination before publication.
Understand the limits
Prompt Lab observations do not reproduce every public chat experience. Avoid claims of probability, causation or broad model behavior that exceed the saved sample.
Quick reference
| Evidence or state | How to use it |
|---|---|
| Prompt version | Create a new version instead of silently editing history. |
| Successful answer | Retain the full response and returned model. |
| Truncated answer | Exclude it from affected denominators. |
| Comparison run | Match settings and repeat count before comparing. |
Check the result
After the reviewed change is deployed, collect fresh evidence with a compatible configuration. A task marked complete records workflow progress; it does not prove the live page changed. Save the verification time and result, then use Verify SEO Fixes for supported checks or run a complete follow-up audit.
Continue your workflow
Return to the documentation home, learn how to interpret audit results, or organize reviewed work in the Action Queue and Fix Generator. Remove private information before sharing exports, and contact LLMIC support with a sanitized example when a result cannot be explained.
Frequently asked questions
Use LLMIC Prompt Lab to design repeatable AI tests with fixed questions, provider settings, evidence and review criteria.
Load the relevant crawl, define the prompt or question set, and connect only the supported provider needed for the measurement. Record the provider, model, location, language, and run time so the observation can be interpreted later.
Review the exact prompt, returned answer, mention or citation, cited URL, provider metadata, and any unsuccessful check. Keep measured provider output separate from local readiness checks and editorial recommendations.
Turn the observation into a review task with the affected page, captured answer evidence, intended audience need, proposed change, and owner. A missing mention is an observation from that test, not proof that the brand is invisible everywhere.
Repeat the same saved question set with comparable provider settings and record the new dated observation. Do not compare different prompts, models, markets, or time periods as if they were the same measurement.
A provider run is a dated observation, not a probability or ranking guarantee. Results can change by model, prompt, location, personalization, availability, and time, and a local checklist cannot prove inclusion in an AI answer.
Save or export the evidence, assign the approved action, and use the related guides linked on the page for the next check. After implementation, repeat the same workflow against a fresh crawl so the result is comparable.