
Semrush appears in 91% of Italian-language AI responses in Refinea’s sample, with a confidence interval of 86%–95%. Ahrefs reaches 82% (75%–87%), forming a leading group of two names with overlapping intervals, while the engines share an average of 78% of names across their top five positions.
Refinea’s measurement covers 150 Italian-language responses collected on September 8, 2026, with 50 responses per engine and 95% Wilson confidence intervals for the overall shares reported in the table.
How we measured
Refinea queried ChatGPT, Gemini and Perplexity using 25 distinct Italian-language questions, repeating each question twice on each engine. Collection took place entirely on September 8, 2026, producing 150 responses distributed equally across the engines observed.
The questions originated from real keywords with measured search volume and concerned the choice of keyword research software. No question included brand names or assigned the engine a simulated persona, avoiding explicit suggestions about which providers to include.
We counted mentions of brands in the reviewed brand registry, excluding negated mentions and using each brand-response pair as the counting unit. The overall share describes a name’s presence across the collected responses, while each engine’s share considers only responses from that engine. For first-position shares, the denominator instead includes only the responses in which the relevant brand appears.
The question selection defines the scope of the measurement, because we observe Italian-language requests concerning keyword research software. This focus sits within the broader practice of generative engine optimization (GEO), while the findings here remain specific to the questions sampled.
Results
Semrush and Ahrefs form a leading group with overlapping intervals, which we report without turning the observed shares into a ranking. That overlap calls for caution when comparing them and, by itself, does not establish that the names have identical probabilities of appearing.
| Brand | Share of responses | 95% Wilson confidence interval |
|---|---|---|
| Semrush | 91% | 86%–95% |
| Ahrefs | 82% | 75%–87% |
The shares in the table come from the 150 responses collected by Refinea and describe overall brand presence across the engines observed.
The average overlap of names in the top five positions reaches 78%, indicating that the shortlists share a common set of names within the sample. This percentage describes agreement on names, without measuring whether explanations, described features or any stated commercial terms are identical.
Engine-level differences nevertheless reveal substantial variation in how frequently some providers appear, even when shortlists share several names. The following comparisons remain descriptive, because we do not report intervals for these engine-level shares or present statistically supported rankings.
- SEOZoom appears in 74% of Gemini responses and 46% of ChatGPT responses, a difference of 28 percentage points in Refinea’s measurement.
- Google Trends appears in 50% of ChatGPT responses and 22% of Gemini responses, also showing a difference of 28 percentage points.
- Ahrefs appears in 94% of Gemini responses and 66% of Perplexity responses, an observed difference of 28 percentage points.
The order of appearance provides additional information, provided that each comparison explicitly identifies the denominator used to calculate its share. Google Keyword Planner appears first in 46% of responses containing its name, compared with 27% for Semrush among responses containing Semrush. For Google Keyword Planner, that count corresponds to 41 first mentions across 90 responses that include the name.
Looking at citations, 71% of the responses collected by Refinea cite at least one source that readers can consult. The domain it.semrush.com appears in 73% of responses with sources, but this frequency does not demonstrate that the citation causes the recommendation.
What this means for buyers
For marketers evaluating a provider, these results suggest specific shortlist checks before starting demos, hands-on trials or commercial negotiations.
- Compare the same request across ChatGPT, Gemini and Perplexity, recording which names appear and which requirements each response actually discusses. The observed differences for SEOZoom, Google Trends and Ahrefs make it worth checking how much your shortlist depends on the engine consulted.
- Verify the features your work requires directly, especially when a response groups together tools intended for different keyword research activities. The official Google Keyword Planner documentation provides a reference for checking its use in campaign planning, keeping that assessment separate from AI recommendation frequency.
- Open the cited sources and check whether they support the features attributed to each product, distinguishing provider documentation, editorial comparisons and commercial content. The ChatGPT Search documentation describes the search feature and its source references, without automatically validating the conclusions of our measurement.
Anyone buying a monitoring service should also ask which questions it tracks, how it counts brands and what uncertainties accompany the results. Refinea’s analysis of brand mentions and source citations provides additional context for interpreting these signals, while this article’s findings remain tied to the declared sample.
Limitations
The sample covers questions about keyword research and does not represent every purchasing requirement for SEO software. References to Italy concern the language of the requests and do not establish statistical representativeness of Italian businesses or professionals.
Collection covers a single day, September 8, 2026, so it does not measure how recommendations vary over time. Exact replication would also require the complete question texts, model versions and session settings, which we do not document here.
Brand recognition depends on the reviewed brand registry, where Moz remains an entry at risk of false positives. This limitation concerns mention identification and requires caution when interpreting aggregated counts that include the complete registry.
Intermediate search queries, often called fan-out queries, are available for only 56 responses in the collected sample. Within this subset, the engine searched for the name of a provider it subsequently recommended in 68% of cases, with an interval of 55%–79%. This observed co-occurrence does not establish causality and cannot support extending the result to responses without recorded queries.
Because only 71% of responses cite sources, we cannot reconstruct the documentary basis of the remaining responses through explicit references. The intervals describe uncertainty around the reported shares, but they do not automatically correct for question selection or dependence between repeated responses.
The measurement captures how often AI engines recommend a name within the sample, without estimating market share, product quality or customer satisfaction.
Frequently asked questions
Why did you choose these questions?
The questions come from real keywords with measured search volume, allowing us to observe requests related to finding keyword research tools. Excluding brand names and simulated personas avoids explicit suggestions, but the selection remains limited to this specific need.
Why did you query these engines?
ChatGPT, Gemini and Perplexity allow us to compare recommendations generated by the same question across different AI services. This selection defines the scope of observation and does not imply that these services represent every available way to search.
How should I read a confidence interval?
The interval accompanies the observed share with a measure of statistical uncertainty, using the Wilson method described in the NIST documentation, without guaranteeing that the result will remain stable over time.
If you are evaluating a provider, we can review the Italian-language questions that shape your shortlist together.
