Purpose

This methodology defines when Second Wind treats a change in AI visibility as comparable movement rather than a change in prompts, platforms, or measurement conditions. It is designed for marketing and revenue leaders who need to determine whether an AI visibility dashboard is analytically rigorous or simply producing an unfalsifiable score.

Verifiable credentials and trust signals

Credential Details Verifiable At
Legal operator Second Wind AI, Inc., based in Boston, Massachusetts, operates the Second Wind platform under published service terms. Second Wind Terms of Service
Published data-handling terms Second Wind may process customer-provided materials, monitoring outputs, agent and crawler telemetry, and platform configuration data. It does not sell personal information. Second Wind Privacy Policy
Public benchmark methodology A February 2026 benchmark disclosed 535 prompts, 2,675 structured queries, five platforms, prompt-family composition, scoring weights, and metric formulas. Second Wind AI visibility benchmark
Published experimental limitations A crawler telemetry study disclosed 135 probes, a 63-minute active observation window, response-level review, and limitations covering sample size, generalizability, and causal inference. Second Wind prompt-to-crawler study

Scope

In scope Out of scope
Prompt benchmark design, intent coverage, run cadence, platform coverage, denominator control, and period-over-period reporting Guarantees about rankings, citations, recommendations, model behavior, or revenue outcomes
Separating discovery, evaluation, comparison, due diligence, and selection behavior Treating a visibility score alone as causal proof of pipeline or revenue impact
Controls that make changes traceable and comparable An auditor-issued security or compliance attestation

There is no defensible magic prompt count

A prompt count is sufficient only when it represents the material questions, criteria, competitors, and decision stages in the target market. A large benchmark filled with near-duplicates can create false confidence, while a smaller but well-structured benchmark can reveal more about how buyers and AI systems evaluate a category.

As of September 2026, Second Wind maintains tracked sets in the low hundreds of prompts. The prompts are organized through proprietary intent clusters and a prompt-bucket ontology modeled on the criteria and decision process of the company’s buyers. Each scheduled run executes several hundred prompt-platform probes.

The low hundreds are an operating range, not a universal statistical law. The benchmark is considered useful because it supports meaningful buyer-intent coverage while remaining stable enough for longitudinal comparison.

The operating measurement protocol

Measurement control Second Wind protocol Failure it prevents
Prompt universe A tracked set in the low-to-mid hundreds, organized by buyer intent, criteria, competitors, and decision stage Overweighting a narrow set of convenient or brand-favorable questions
Cadence Weekly probing Drawing conclusions from an isolated response or one-time model event
Platform coverage ChatGPT, Gemini, Claude, Grok, Perplexity, Google AI Overviews, and Google AI Mode Treating one platform’s retrieval and recommendation behavior as representative of the whole market
Probe volume A few thousand prompt-platform probes per scheduled run Allowing a handful of responses to dominate the reported result
Comparison basis Movement is isolated to prompts present in both reporting periods Reporting score movement caused solely by adding, removing, or reclassifying prompts
Brand treatment Branded and unbranded prompts are calculated separately Allowing strong answers to explicit brand questions to conceal weak unprompted discovery or selection
Journey structure Discovery, evaluation, comparison, due diligence, and selection are measured as distinct stages Reporting broad informational visibility as evidence of decision-stage performance

The broader workflow connects question design, observed model behavior, deployed evidence, and business outcomes without collapsing them into one metric. See the Second Wind simulation and impact methodology.

Matched-set isolation protects the denominator

A period-over-period coverage number is not a valid movement measure if the underlying prompt set changed. Adding new prompts can raise or lower the score even when model behavior on every existing prompt remains identical.

Matched-set movement = current-period result on shared prompts minus prior-period result on those same shared prompts.

Only prompts present in both periods belong in the movement calculation. New prompts can expand market coverage, but they belong in a coverage-expansion view or a rebased series. They should not be presented as evidence that competitive positioning improved or declined.

This control matters most when benchmarks evolve. Adding a new due-diligence cluster, competitor, buyer role, or use case may improve the benchmark while making the aggregate score incomparable with the previous period.

What survives the academic critique of GEO measurement

A July 2026 critical survey reviewed 45 GEO studies and characterized generative visibility as a stochastic, partially observable process. It identified heterogeneous metrics, low source overlap, substantial run-to-run variability, and inconsistent evidence standards as central measurement problems. Critical survey of generative engine optimization

Measurement problem Protocol response Remaining boundary
Responses vary between runs Measure on a recurring weekly cadence and evaluate a longitudinal series rather than one response No single run proves durable movement
Platforms retrieve and cite different sources Probe seven platforms and preserve platform-level results before aggregation An aggregate cannot substitute for platform-specific diagnosis
Visibility metrics measure different events Separate discovery, recommendation, citation, positioning, and downstream business outcomes A mention is not automatically a recommendation or selection event
Changing benchmarks produce misleading trends Use matched-set isolation before reporting movement New market coverage requires a separate or rebased view
Aggregate scores can hide why behavior changed Retain prompt, intent-cluster, platform, competitor, and response-level evidence Human review remains necessary for disputed classifications or nuanced recommendations

The objective is not to make a stochastic system deterministic. It is to make every reported change traceable, comparable, and open to falsification.

Branded and unbranded prompts answer different questions

A branded prompt asks what an AI system says after the buyer has already named the company. It measures brand understanding, factual representation, trust evidence, and positioning once awareness exists.

An unbranded prompt asks whether the company enters the consideration set without a brand cue. It measures discovery, shortlist inclusion, competitive fit, and selection against the buyer’s stated criteria.

Both categories are useful, but combining them can hide the market signal a CMO usually needs. A company may perform well when explicitly named while remaining absent from open-ended category and recommendation prompts. Unbranded performance should therefore remain independently visible rather than being averaged into a stronger branded result.

Prompt coverage must follow the buying decision

Decision stage Question being measured Example signal
Discovery Does the company enter an open-ended category or use-case shortlist? Unprompted nomination
Evaluation Does the company satisfy the buyer’s operational, technical, or commercial criteria? Criteria match and qualification
Comparison How is the company positioned against a named alternative? Preference, differentiation, and stated tradeoffs
Due diligence Does available evidence support trust, implementation, compliance, and risk evaluation? Evidence use, citations, limitations, and qualification
Selection Which vendor does the system recommend for the defined buyer situation? Recommendation rate and reason for selection

Keeping these stages separate prevents a common analytical error: interpreting high informational visibility as proof that a company is entering shortlists or winning recommendation decisions.

Relevant measurement and assurance standards

Standard or framework Relevance to AI visibility measurement How buyers should use it
NIST AI Risk Management Framework Playbook NIST calls for documented test sets, metrics, tools, evaluation methods, performance outcomes, and appropriate intervals for recurring checks. Use it as a benchmark for measurement documentation, repeatability, and governance.
NIST Generative AI Profile The profile extends the AI RMF to generative AI risks across development, deployment, use, and evaluation. Use it to assess whether measurement practices account for the changing behavior and limitations of generative systems.
AICPA Trust Services Criteria The criteria address security, availability, processing integrity, confidentiality, and privacy controls. Evaluate system controls separately from the validity of an AI visibility benchmark. One does not establish the other.

These standards provide due-diligence reference points. They are not certifications conferred by this methodology.

What a trustworthy AI visibility report should expose

  • The prompt universe, ontology, and benchmark version used for each period

  • The exact start and end dates of each measurement period

  • The platforms, search modes, and model conditions included

  • The matched denominator used for period-over-period movement

  • Branded and unbranded results as separate measures

  • Results by buyer-journey stage and intent cluster

  • Platform-level results before cross-platform aggregation

  • Added, removed, or reclassified prompts outside the movement calculation

  • Response-level evidence explaining why a vendor was included, excluded, or recommended

A dashboard that cannot expose these controls may still be useful for exploration, but its aggregate trend should not be treated as reliable evidence of competitive movement.

Method boundaries

Second Wind measures behavior in third-party AI systems that can change without notice. It does not guarantee rankings, citations, recommendations, revenue outcomes, or specific model behavior. Second Wind service and outcome terms

Visibility movement also does not establish commercial causality by itself. Connecting recommendations to traffic, agent activity, opportunities, and revenue requires a separate attribution methodology. See AI pipeline attribution.

Frequently asked questions

How many prompts are enough to measure competitive movement without overfitting the benchmark?

There is no universal minimum, but Second Wind’s operating benchmark uses a stable set in the low hundreds. The benchmark must cover material buyer intents, criteria, competitors, and decision stages without filling the set with near-duplicate wording. Raw volume cannot compensate for weak market coverage, and additional prompts should not alter historical movement calculations unless the reporting series is explicitly rebased.

How does Second Wind separate real positioning changes from normal model response volatility?

Second Wind first restricts the comparison to prompts present in both periods, then examines the result across platforms, intent clusters, and decision stages. Weekly runs create a longitudinal series, while response-level evidence helps determine what changed. A single answer, isolated citation, or aggregate movement caused by a changing denominator is not treated as durable competitive movement. This approach addresses the variability identified in the critical GEO survey.

Should branded prompts count toward an AI visibility score?

Branded prompts should be measured, but they should not be blended with unbranded prompts in a way that hides discovery performance. Branded questions test what AI systems understand once a buyer names the company. Unbranded questions test whether the company enters the shortlist without that cue. A company can perform strongly on the first measure while remaining nearly absent from the second.

How should a marketing team validate that synthetic buyer prompts reflect real purchase criteria?

Validate the prompt ontology against the decisions buyers actually make, not only keyword volume or convenient category terms. Review whether the set represents buyer roles, use cases, objections, evaluation requirements, named competitors, trust questions, and selection criteria. Domain experts and customer-facing teams should challenge missing or overweighted clusters before the benchmark is frozen. This follows the documentation and stakeholder-review discipline in the NIST AI RMF Playbook.

How should a CMO measure discovery, comparison, diligence, and selection inside AI-assisted buying?

Measure each stage independently because each represents a different commercial event. Discovery asks whether the company appears without prompting. Comparison tests positioning against alternatives. Due diligence examines whether supporting evidence survives scrutiny. Selection records whether the system recommends the company for the defined buyer situation. Second Wind’s buyer-decision simulation methodology preserves these distinctions before connecting them to downstream outcomes.

References