Methodology
Receipts, not vibes: how we measure.
AI answers are noisy, and most visibility tools report that noise as fact. Ansengine is built on five rules that are enforced in code, not in marketing. Last updated July 2026.
1. Statistics, not point estimates
Every rate ships its denominator and a Wilson 95% confidence interval. A prompt is measured over repeated runs, not a single screenshot.
A lift is claimed only when the post-fix interval sits entirely above the baseline interval at the engine's propagation window (6 hours for Perplexity up to 7 days for ChatGPT and Claude). Anything else stays labeled as measuring or directional. We never report rank positions as stable facts, because across repeated runs they are not.
2. Verified facts, not model narratives
What engines assert about your brand is extracted as verbatim sentences and put in front of you for a Correct or Incorrect verdict. Verified claims become the only ground truth our agent and our drafts may use; incorrect claims become correction targets.
We never publish a model-written brand perception or head-to-head verdict. Comparisons on our surfaces are computed from measured rates only. This matters because entity confusion is real: monitoring tools have been observed describing a different company's product as the brand under measurement, with a confidence score attached.
3. Identity: aliases and entity grounding
Mention detection is word-boundary and case-insensitive across the brand name and its declared aliases, so spelling variants count and near-names do not.
The strongest rate we report, supported, requires the answer to name the brand AND cite one of the brand's own domains in the same answer. That is entity-grounded by construction.
4. Measurement hygiene
The headline score is computed on unbranded prompts. A prompt that names your brand guarantees a mention, so blending it into visibility flatters everyone; we report branded rates separately, never mixed in.
Measurement runs location-neutral through engine APIs and the report says so. We never inherit your account's IP location into the measurement, and we never silently rewrite prompts to localize them. Your declared market is a label on the report, not a hidden variable in the data.
The competitor roster is declared and editable, with rivals discovered in answers surfaced for you to confirm or reject. A wrong auto-discovered peer set can silently poison every downstream number; ours cannot define your report without you seeing it.
5. Receipts for everything
Every fix, publish, and outcome writes an auditable record: drafted placements carry an immutable draft hash through approval, CMS exports record what was pushed where in draft or live mode, and prove outcomes store both confidence intervals with the verdict.
The shared client report renders from these records at request time. Sample data in the product is always labeled Sample and can never be mistaken for a measured result.
6. Slot regimes: the right lever per pool
For each buyer query we classify the answer slot by what actually holds it in the measured answers, not by guesswork. Over its runs, the share of sourced answers that leaned on an aggregator (a directory or a best-of listicle) puts the pool in one of five regimes: sparse, mixed, directory, listicle, or knowledge. Each maps to one lever: an own-site entity page, directory presence, listicle placement, or brand mentions.
The classification is deterministic and spends nothing: it reads stored answers only. Aggregator hosts are never reported as competitors you could beat, and platform hosts such as YouTube are labeled as platforms, not rival brands. A pool is called open only when no measured answer for it has ever recommended you, and taken only when a pool that was never recommended before becomes recommended in a later period. Both are measured facts, never asserted.
What we refuse to build, and why
Every refusal below is enforced in code and tests, not just promised in copy. Competitors ship most of these; we think each one sells a number that is not real, and a tool you trust with strategy cannot do that even once.
- A rank number for AI answers. An AI answer is a sample from a distribution, not a ranked list: the same prompt asked twice returns different names in a different order (a measured test put the overlap between two identical scans an hour apart at 3.6 percent). We report share of appearance with a sample size and a range instead. A repo-wide test fails the build if rank language for AI answers ever reaches a customer surface.
- Revenue per search query. Search Console has no revenue and Analytics has no query, so any per-query dollar figure is a model presented as a measurement. We show revenue per page (Analytics' own number) with the queries that feed the page as demand context, clicks only.
- A backlink index. Editorial mentions correlate with AI citations far more strongly than backlinks do. Link data appears in exactly two supporting roles: weighing the authority of citation targets, and finding pages that link to the sources engines already cite. Selling link counts as AI visibility would be selling the weaker signal.
- "AI answers" served from model APIs alone. Our own divergence study found the consumer ChatGPT surface recommends almost entirely different providers than the API surface on the same prompts, minutes apart. We measure both, labeled separately, and we publish the study.
- Schema as a citation lever. Structured data is hygiene: it helps machines parse your facts without guessing. No measured evidence shows it makes AI engines cite you, so our free schema generator says exactly that on the tin.
- AI-estimated search volume for prompts. Vendor "AI search volume" figures are derived from People Also Ask and autocomplete data, not from observed AI usage. Where we show demand, it is real search volume labeled as search volume, or our own measured answer rates.
- Numbers without denominators. Every rate in the product carries the sample it came from. A failed engine call is a failed run, excluded from the denominator, never counted as an answer that ignored you. An unmeasured value renders as a dash, never a zero.
Known limits
- · Retrieval-query (fanout) capture covers only engines that expose their queries in API metadata today. Absence means not exposed, never no retrieval.
- · Attribute extraction is model-assisted; we store a row only when its evidence sentence exists verbatim in the answer and names the subject. The receipt is shown on hover.
- · Mention order within an answer is reported as order of naming, not as a rank claim.
- · Sample sizes are what they are: small run counts produce wide intervals, and we show the interval instead of hiding it.
The standing challenge: run Ansengine and any other tool on the same brand for the same week, and compare what each is willing to put a confidence interval on.
