← back to all posts

Contrast Relevance Analysis: An AI Agent That Compares a Page With the Top of Google

Contrast Relevance Analysis: An AI Agent That Compares a Page With the Top of Google

"Why is the competitor above us?" is a question clients ask on every other call. An honest answer takes a contrast analysis: open the top results for the query, write down what every competitor covers, and compare it with the client's page. That's about three hours per query, so by hand agencies check one or two queries out of the dozens worth checking. The rest gets covered by "relevance analysis" tools that return a single number and never say what to add. This article builds an agent that takes a page URL and a query and returns a list of edits, sorted by how many top-ranking competitors already have each one. And it refuses to suggest edits when there's nothing to compare. The setup was tested on a live store — the results are below.

What this agent should actually do

One job: the client's page and a search query in, a list of edits out. The first line is what nine competitors out of ten have and the client doesn't; the last is what four of them have. Along the way the agent measures five things:

  1. Subtopic coverage. Which questions each top-ranking page answers, merged into a checklist with frequencies.
  2. Facts and entities. Prices, timelines, quantities, names and specs that most of the top states and the client doesn't. The cheapest to count and the most concrete in the output.
  3. Direct answer. Whether the page answers the query itself within the first 500 characters.
  4. Freshness. The median update date in the top versus the client's.
  5. Demand. How many people search for the phrase, to tell an important gap from a gap in a query nobody needs.

What it should not do: output a combined score, promise rankings, suggest edits when the top answers a different question, or invent demand when the tool returned none. Deciding which edits to make, and in what order, stays with the specialist.

Why there's no overall relevance score

The usual tools fold breadth, depth, BM25 and domain authority into one score. The problem is that the ranking model isn't visible from outside, so any such score is a guess about how it works. When we took apart a table from one popular tool, the final sort turned out to be a 0–100 percentage multiplied by a BM25 weight between 0.08 and 0.27, and the rank correlation of that order with real positions was about 0.47 — the order explained roughly a fifth of the variance. And even an accurate score wouldn't say which component dropped.

The framing here is different: the search results are already labelled data. The question isn't "rate my page" but "what do the pages that rank have that mine doesn't". The metric is a contrast — the client against the top for a specific query. Hence the name.

How to phrase the task

The agent needs two things in the task text: the client's page URL and the search query. They're extracted by different mechanisms:

What How it's extracted What that means in practice
Client URL The first link in the task text Put exactly one link in the task.
Search query The first phrase in quotes, then the phrase after the word "query", then the last line Put the query in quotes and make it the first quoted text.

A phrasing that works:

Compare the page https://example.com/seo-tbilisi with the Google top
for the query "seo tbilisi" and tell me what it's missing.
Without quotes, the failure is silent. The query becomes the whole last line of the task, including "tell me what it's missing". A phrase nobody searches for goes to the SERP, random competitors come back, and the report comes out coherent and plausible — a careful comparison against the wrong top.

The router: seven steps

A contrast analysis doesn't fit into a single model call, so this is a router — a role with a numbered algorithm instead of a single instruction. How routers work in general is covered in the guide "How to Create an AI Agent. Site Audit". Here are this one's steps as they appear in the panel:

Steps 1–3 · data collection, run at the same time
1
{clean_task}
Page fetch
Opens the first URL in the task — the client's page. That's why the task must contain one link.
2
{clean_task}
Google SERP for a query
Positions, domains and full URLs of competitor pages — what the collector downloads next.
3
{clean_task}
Bing demand for a query
or Wordstat for Russian queries
Choose by the language of the project's queries. Wordstat only knows Yandex's Russian-language base and returns zero for an English phrase; Bing works in Russian too, but on its own, smaller base.
Steps 4–7 · subagents
4
contrast_page_collector || {current_context}
Call subagent
The collector doesn't need the original task: the competitor URLs are already in the context after step 2.
5
contrast_corpus_analyst || {current_context}
Call subagent
6
contrast_gap_auditor || DATA FROM PREVIOUS STEPS: {current_context} {db_data_context} Query and client URL: {clean_task}
Call subagent
The data comes as a block before the task, so the model sees the numbers before the instruction not to invent them.
7
contrast_reporter || DATA FROM PREVIOUS STEPS: {current_context} {db_data_context} Original task: {clean_task}
Call subagent
The last step's answer is the agent's answer — the router returns it to the user in full.

Three consecutive tool steps run at the same time: the client's page, the SERP and the demand data arrive in parallel, not in a queue. The router's system prompt does almost nothing — its job is to carry the URL and the query to the last step intact:

You orchestrate a contrast analysis of a page. Your only job is to
carry two things through every step without distortion: the client's
page URL and the search query the comparison is made for.

The analysis itself is done by tool steps and subagents. Do not write
conclusions of your own, do not retell step results and do not add
your own considerations to them.

If the task doesn't contain a page URL or a search query, say so right
away and don't run the steps: without either one, the comparison is
meaningless.
The role key must contain a _router segment, and background mode is mandatory. Without the segment the platform doesn't treat the role as a router: the per-URL cache is lost, and instead of a report the user gets a raw "STEP 1… STEP 7" log. The segment is matched whole, between underscores: contrast_analysis_router works. And seven steps that download a dozen pages take minutes, so turn on "Always run in background": the agent acknowledges the task right away and sends the report as a separate message.

Four subagents, one job each

You could write one big instruction: "compare the page with the top". In practice a model given both the check and the recommendations switches into advice mode halfway through the table and stops checking the remaining rows. So the work is split across four roles.

Top-page collector

A regular tool step can only open the first URL in the task — the client's page. Only a subagent that calls "Page fetch" itself can walk through ten different URLs. It also compresses each page into a summary right away: ten full texts simply won't fit further down the chain.

You collect competitor pages in a contrast-analysis pipeline. The
context contains the Google SERP for a query: positions, domains and
page URLs.

For EACH result, call fetch_url_content on the page URL (not the
domain) and work out what's on it.

Do not retell pages in full or quote long passages. The next subagent
reads your answer, and ten full texts simply won't fit. For each page
return exactly this block:

POSITION N — URL
Title: the page's H1
Type: commercial / informational / catalogue / product card / other
Answers: 5–12 subtopics, one line each, as specific as possible —
  not "about prices" but "how much it costs and what's included";
  not "about timelines" but "how many months until results". Phrase
  them as a customer's question, not a site section.
Facts and numbers: specifics the page states — prices, timelines,
  quantities, tool names, guarantees. Verbatim, with figures. If there
  are none, write "no facts".
Direct answer: yes / no — does the page answer the query itself in
  the first 500 characters, before navigation and intros.
Length: approximate word count of the main text.
Updated: the date, if the page shows one; otherwise "not stated".

If a page didn't open or returned anti-scraping protection, write
"POSITION N — URL — DID NOT OPEN" and move on. Don't invent the
content and don't guess it from the SERP snippet: an invented page
spoils the whole checklist downstream.

Finish with the line "Pages collected: X of Y".

Top-results checklist

Merges the summaries into one list: which subtopics and facts appear on how many competitor pages. Calls no tools. This is also where the first intent check lives — whether the top is taken by a different entity with the same name.

You are the corpus analyst in a contrast-analysis pipeline. The context
contains summaries of the pages in Google's top results for one query.

Your job is to merge them into a single checklist for the topic.

1. MERGE SUBTOPICS. Collect every "Answers" line from every page.
   Merge wordings that mean the same thing — "how much it costs" and
   "service prices" are one subtopic, not two. For each, count how
   many of the collected pages have it.

2. MERGE FACTS. The same for numbers and specifics: which figures,
   timelines, guarantees and names most pages state. Flag separately
   the facts repeated by three or more competitors — that's the
   industry-standard answer, not someone's find.

3. PAGE TYPE AND INTENT. Which page type dominates the top, and by
   how much (e.g. "7 of 9 are commercial landing pages"). That signals
   what search considers a fitting answer to the query.

   Separately, check whether the top is taken by a DIFFERENT entity
   with the same name. Signs: artist, company or film cards, streaming
   profiles, reference and navigational pages instead of topical ones.
   If so, output the line "TOP INTENT: <what's actually there>"
   and don't force those pages' subtopics onto the query's topic.

4. NORMS. Median text length and median update date across the
   collected pages.

Count ONLY the pages actually collected. If the collector marked some
pages as not opened, the denominator shrinks, and you must say so
explicitly: "counted over 7 pages of 10". Don't substitute snippets
for them and don't fill in the picture from the model's memory.

Output format:
— Table "Subtopic → how many of N → industry standard (yes if 3 or
  more)", sorted by frequency, descending.
— Table "Fact/figure → how many of N → example wordings".
— "Top norms" block: dominant page type, median length, median
  update date.
— Line "Counted over X pages of Y".

Do not output any combined scores, relevance ratings or page rankings.
Your result is a count per subtopic, not a single number.

Client check

Walks the checklist row by row and assigns one of three verdicts. The client's page is already in the context from step one, so the auditor doesn't go back to the site. It writes no recommendations — that's the next role's job.

You are the auditor in a contrast-analysis pipeline. The context
contains the top-results checklist (subtopics and facts with counts)
and the client's page text collected in the pipeline's first step.
Do not fetch the site again.

FIRST — DOES THE INTENT MATCH. Determine the client's page type and
compare it with the dominant page type and intent from the checklist.
If they're about different things — a tutorial versus artist cards, a
commercial landing page versus a reference site, a product versus a
forum — there's nothing to compare. Write "INTENT MISMATCH: <what the
top is about> versus <what the client's page is>", explain in two
lines why the page ended up in these results at all, and stop there —
don't go through the checklist, don't assign verdicts.

If the intent matches, go through the checklist STRICTLY row by row,
skipping nothing, and give each subtopic one of three verdicts:

  PRESENT — covered, no worse than competitors;
  WEAK — mentioned, but without the specifics the top has (e.g. the
    price is "on request" while seven of ten give a range);
  MISSING — not covered at all.

Back each verdict with a quote from the client's page or an explicit
note that it's absent. Don't assume: if the page has a wording but it's
ambiguous, that's WEAK, not PRESENT.

Check four things separately:
— Direct answer: does the client's page answer the query itself in
  the first 500 characters. How many competitors do — see checklist.
— Facts: which figures and specifics from the "industry standard" are
  missing from the page. List them by name.
— Length: the client's word count versus the top's median. Word count
  means nothing by itself — give it only next to how many subtopics
  are not covered.
— Freshness: the client's update date versus the top's median.

Output format:
— Table "Subtopic → how many of N in the top → verdict → evidence".
— "Direct answer / Facts / Length / Freshness" block — one line each.
— Nothing else: the next step writes recommendations, not you.

Report with edits

Turns the verdict table into the whole point of the exercise — a list of edits by descending frequency. And checks two conditions under which there will be no list at all.

You write the contrast-analysis report. The context contains the
top-results checklist, the result of checking the client's page and
demand data.

Build a report for the person who will edit the page from it.

TWO CONDITIONS WITHOUT WHICH THERE IS NO LIST OF EDITS. Check them
before writing the first line.

1. INTENT. If the check says "INTENT MISMATCH", do NOT write a list of
   edits at all. The report is one explanation: what the top is about
   for this query, why the client's page doesn't belong to it, and
   which queries to target instead. Never force another topic's
   subtopics onto the page, even if formally "all competitors have
   them".

2. SAMPLE SIZE. If FEWER THAN FOUR top pages were collected, there is
   no list of edits either. "2 of 2" looks like consensus but actually
   means there's nothing to compare against. Say how many pages
   opened and why that's too few, and suggest re-running later. With
   four or five pages collected you may write the list, but warn in
   the report's first line that the sample is small and the
   frequencies are rough.

THE ORDER OF EDITS is the heart of the report. Sort by the number of
competitors that have it, from most to fewest. A gap nine of ten have
closed goes above one four have closed. Within the same frequency,
cheaper work goes higher.

Each edit gets exactly four lines:
  What to add or rewrite (specifically — not "improve the pricing
    section" but "state the price range and what it includes").
  How many top competitors have already done it.
  What exactly they write — 1–2 example wordings from the checklist.
  Why it's here: a quote from the check or the note "not on the page".

Then a short "Already good" block — subtopics marked PRESENT, one line.
It keeps the edits from reading like a verdict on the page.

Then a "Query context" block: search volume from the demand data and
refined queries worth covering on the same page.

If there is no demand data for this phrase, say exactly that: "no
demand data for the query". Do NOT invent a list of related queries:
without a source they're guesses, and in a report they look identical
to facts. For long-tail queries, missing data is the norm.

The ban is wider than query lists: NO claims about demand, seasonality
or user behaviour without a number returned by a tool. If there's a
number, give the number; if there's no number, say nothing about it.

Before writing "no data", request demand for the topic's head phrase —
a shorter, broader one. If that has data, give it, clearly marked as
demand for the topic rather than for this specific query.

WHAT YOU MUST NOT DO:
— output a combined score, relevance rating or percentage match with
  the top: one number hides which edit matters;
— promise rankings: the analysis shows a content gap, not the reason
  for a ranking;
— count pages that didn't open — if the check ran on an incomplete
  set, put that in the report's first line;
— if SERP or demand data didn't arrive at all, no report is written.
  Say plainly that the source didn't respond and that no conclusions
  can be drawn from this run. An empty SERP almost never means "no
  competitors".

Format: a heading, the list of edits by descending frequency, "Already
good", "Query context". No intros about the importance of SEO.

Tools and integrations

Everything the router needs already ships with Orakul and needs no separate setup: "Page fetch", "Google SERP for a query", "Search volume (Wordstat)", "Bing demand for a query" and "Call subagent" are picked from the step's dropdown. Bing needs a Bing Webmaster Tools key, which is issued to any verified account with no ad spend.

Google and Yandex SERPs, Wordstat and Bing demand are Agency-plan tools. On Pilot and Team the agent sees data about the client's site but not the market, and a contrast analysis can't be built.

A live run: the most useful result was a refusal

The agent was run on a real store — anastasiya.studio, which sells printable music worksheets for kids. The full write-up with numbers is in the article "Which Page-Two Rankings Are Worth Fixing", and the agent's report on all five pages is in a public document. Here's what matters for the setup.

Candidates were picked by another agent from Search Console data for August 13 – September 12, 2026: pages with clicks at positions 10–20. Clicks confirm there's demand, and page one isn't far away.

Query Position What the analysis showed
music note tracing worksheets 10.0 Top held by big players; a ceiling of 1–2 positions
music color by number 10.5 A clear list of edits
music coloring sheets for elementary free 15.5 A clear list of edits
bass clef note naming worksheets pdf 16.2 Intent mismatch — no edits
twinkle twinkle little star game 18.4 A game page with no text, nothing to compare

For "music color by number" the report looks the way it was designed to: worksheet count on the product card — 4 competitors of 6; an explanation of the mechanics and color key — 4 of 6; ages and grade levels — 4 of 6; file format and how to download — 4 of 6; skill benefits in visible copy — 3 of 6.

But the fourth query turned out the most useful. Instead of a checklist, the agent wrote that there was nothing to compare: the results are direct PDFs from school sites and free directories, while the store's page is a $2.50 product with a cart. The searcher wants a free file, and added subtopics won't change that. On top of that, only 3 of the 9 results returned HTML. The agent offered two paths: retarget the product page to commercial queries, or build a separate free page with a couple of PDFs that links to the paid pack.

Three things changed on the site right after the report: "A4 PDF" and "Instant Download" badges on every product, a "grade level" field on the product card, and free products handing over the file immediately. Snippets were deliberately left alone: snippets move CTR, content moves position, and change both at once and a month later there's no telling which one worked.

Three safeguards that came from real failures

The prompts above weren't written in one go. Every "stop" condition in them traces back to a specific run where the agent produced a neat and useless report.

  • Intent. For "music trace", Google's entire top was taken by cards for a DJ of that name — Apple Music, Insomniac, streaming profiles. The agent dutifully compared a blog post about note-tracing sheets for kids against them and suggested adding the House genre, a discography and club dates in Tashkent. The frequencies were counted correctly; only the name matched. Now the analyst looks for a "different entity" in the top, the auditor stops on a mismatch, and the report writer explains instead of listing edits.
  • Sample size. In the same run, two pages out of six opened, and "2 of 2" looked like competitor consensus. The usual non-openers are Pinterest, Scribd and direct PDFs — they don't return HTML. Below four collected pages, no list of edits is produced.
  • No invented demand. Wordstat returns zero for an English phrase, and Bing answers for broad phrases and stays silent on the long tail: "coloring pages" got 46,495 impressions over four weeks, "music note tracing" got nothing. On early runs the model, left without data, wrote in plausible related queries, and in the report they looked like data. Now: if there's a number, the agent gives the number; if not, it writes "no data".

One more rule from the same series: an empty SERP is a source failure, not a free niche. The SERP and demand tools distinguish "nothing found" from "the source didn't respond" and say so plainly, and in that case the report writer doesn't write a report.

How to pick pages and check the result

Running the analysis on every page is expensive and pointless. The working scheme from the store:

  1. Candidates are pages with clicks at positions 10–20 from Search Console, with an impression floor. Without one, flukes get through: on a manual check, one page + query pair at position 10 rested on 2 impressions and 1 click on a single day.
  2. Check one row of the export by hand in the Search Console interface.
  3. Change one lever at a time — content separately from snippets.
  4. Check after 4–6 weeks: the average position of the page as a whole, not of page + query pairs — with two impressions any change is meaningless. And not clicks — with seasonal demand they'll drop whatever you change.
  5. With seasonal demand, repeat on a schedule. The store needs a fresh export every two weeks, and seasonal pages are better analysed before the season: edits take weeks to reach the results.

So the one-month check doesn't depend on a specialist's memory, turn on "Track outcome" on the router with a 720-hour interval. The agent saves its recommendations under the page URL and a month later asks the person who set the task, in Telegram, whether the edits were made. The history is keyed by URL, so it also collects the results of other agents that worked on the same page. More in the guide "How to Create an AI Agent. Account Manager".

The knowledge base for this role

This router barely needs a knowledge base: everything it says comes from the SERP, competitor pages and demand data. The client's constraints — "we don't publish prices", "we don't promise timelines" — are more reliable written straight into the task text, after the quoted query: step 7 passes the original task to the report writer in full, so it won't suggest an edit the client can't make.

The knowledge base comes in at the next stage, when the list of edits turns into copy. The copywriter writes the missing blocks in the client's voice if you name that client's knowledge base in the request. Build it from the client's already-approved pages, in one .md file, not ten.

Testing and tuning

Do the first run on a page of your own where you know exactly who ranks and why. Compare the report with what you'd have done by hand. Don't look for errors in the frequencies — the model counts them from the collected summaries — look for a subtopic phrased too broadly ("about prices" instead of "how much it costs and what's included"). That's fixed in the collector's prompt, not the report writer's.

Router step results are cached for 24 hours per page URL. You can rewrite the report writer's prompt and re-run it, paying only for the last step. But the cache has a flip side — see the first item below.

Common mistakes

  • A second query for the same page within a day. The router cache is tied to the page URL, not to the query. Check a page for "seo tbilisi", then an hour later for "website promotion tbilisi" — the second report is built from the first query's SERP and checklist. Space different queries for one page a day apart.
  • A query without quotes. The last line of the task becomes the search phrase, and the report carefully compares the page against the wrong top.
  • Two links in the task. Step one opens whichever comes first, and that's the page compared with the top. A competitor link "for reference" breaks the whole run.
  • Taking the page's top query without checking it. A page can rank for a phrase it has nothing to do with, like the note-tracing blog post for "music trace". Look at the query yourself before handing it to the agent.
  • Changing everything at once. Content, snippets and internal links in one week — and a month later nobody knows what worked. One lever at a time.
  • Adding an "overall score" to the report. The moment a combined score appears, the report stops saying what exactly to fix.

 

≈ 3 hours

≈ $300 saved per month at 10 clients

About three hours per query by hand — opening the top ten, writing down what each competitor covers, checking the client's page and writing up the edits. At two queries per client per month (the two-week cycle from the store above) and 10 clients, that is 60 hours; at the fully-loaded junior rate this site uses throughout ($850/mo ÷ 160 h ≈ $5/h), about $300 a month. Substitute your own rate and your real hours — the result comes out far higher.

Максим Сафьянов
Максим Сафьянов

I build Orakul: a self-hosted AI-agent orchestrator for digital and SEO agencies — unlimited roles, your own server, source code handed over at setup.

What is contrast relevance analysis?
It's a comparison of a page with the pages already in the top results for a specific query. Instead of a combined score it gives a list of subtopics and facts that most competitors have and the client lacks, sorted by how many competitors have each one.
How is contrast analysis different from tools with a relevance score?
Those tools fold different metrics into one score, even though the ranking model isn't visible from outside and a single score doesn't show what to fix. Contrast analysis sums nothing up: every line of the report is a specific edit with the number of competitors that have already made it.
Which pages are worth a contrast analysis?
The best candidates are pages with clicks at positions 10–20 in Search Console, with an impression floor: demand is confirmed and page one is close. Before the analysis, check that the query actually matches the page’s topic.
What does the agent do if the top is a different type of page?
It stops. If the top answers a different question — for example, the results are free PDFs while the client has a paid product page — the agent doesn't write a list of edits; it explains the intent mismatch and suggests which queries to target instead.
Does contrast analysis require the Agency plan?
Yes. Google and Yandex SERPs, Wordstat and Bing demand are part of the Agency plan. On other plans the agent sees data about the client’s site but not the market.

No comments yet — be the first to share your thoughts.