Understanding Your Metrics

Branded prompts: answer score and themes

Branded prompts ask about you directly. Their report answers a different question from the rest of the product: not "are we visible", but "is what the AI says about us accurate and fair".

Every number here is measured over runs where the AI judge successfully returned a verdict.

The four cards

Sentiment

How positively AI describes you when asked directly, from -1.00 to +1.00. It is the average of the judge's tone score across all judged branded runs.

Answer score

The quality of the AI's answer about you, blending factual accuracy, completeness and clarity into a single 0 to 100 figure. The judge is told to score conservatively: 100 is perfect, around 60 is adequate, 30 or below is poor or wrong.

Competitor co-mention

How often the AI brings up a rival when it was only asked about you.

                    judged runs naming at least one rival
Co-mention rate =   -------------------------------------  x 100
                              all judged runs

A run counts once no matter how many rivals it names. This card's arrows are inverted: a falling co-mention rate shows green, because being compared less often when someone asks about you directly is a good outcome.

Owned citation

How often the AI cites your own website when answering a question about you.

                       judged runs citing your domain
Owned citation rate =  -------------------------------  x 100
                              all judged runs

A low number here usually means your own documentation is not structured for AI to quote. That is what the GEO score is for.


Sentiment over time

Sentiment over time, with a thick average line and one thin line per engine
Sentiment over time, with a thick average line and one thin line per engine

The thick line is the overall daily average; the thin dashed lines are individual engines. The axis runs from -1 to +1 with a marker at zero.


Top themes

The most useful section on this report, and the one that most directly shapes messaging. The judge records the recurring ideas it used to describe you, and whether each is framed positively or negatively.

ThemeSentimentPresence
Client reporting+0.7754%
Custom workflows+0.8638%
Beginner suitability-0.4731%
Cost effectiveness+0.3331%
Progress tracking+0.6015%

Presence is how often the theme actually appears:

            answers that actually stated this theme
Presence =  --------------------------------------  x 100
                    all judged branded runs

The denominator is all judged runs, not only the runs that raised the theme.

How to read the example above. Everything about this brand is framed positively except one row. "Beginner suitability" sits at -0.47 and comes up in nearly a third of answers. That single line identifies the perception problem precisely: the AI thinks the product is excellent, but not for newcomers. That is a concrete content gap, not a vague sentiment number.


Uninvited competitor co-mentions

Which specific rivals the AI volunteers when nobody asked.

            times this brand was named across all answers
Rate =    ----------------------------------------------  x 100
                      all judged branded runs

These names come from the judge reading the answer, so they include brands you are not tracking. This list is often where you discover that the rival the AI most associates with you is not the one you have been benchmarking against.


Accuracy flags and missing facts

Places where the AI may be wrong, stale, or leaving out something important.

KindMeaning
errorThe AI stated something factually wrong
warnThe information may be out of date or contested
missingA key fact is absent from the answer

A count appears when more than one answer raised the same point, which is the signal that it is systemic rather than a one-off.

This is the most directly actionable section in CitedSpy. If two separate flags say your pricing looks outdated, the fix is to make current pricing more prominent and machine-readable on your own site.


Frequently asked questions

They measure different things. Answer score rates how accurate and complete the answer is; sentiment rates how warmly it describes you. A thorough, fair, slightly critical answer scores high on quality and moderate on tone.

Usually yes. If AI repeatedly volunteers a company alongside you, it is a competitor in the eyes of the model, whether or not it is on your list.

Look at the count. One flag on one answer may just be that engine's phrasing. The same flag on several answers is a real gap in what the web says about you.

Need more help?

Our support team is here to help you get the most out of CitedSpy.

Contact support