AigeoRadar

AigeoRadar · AI Context Benchmark

AI Context Benchmark

Question → sufficient layer(s) → retrieved slice → measured cost → evidence

Demo — AigeoRadar Pricing

https://aigeoradar.com/pricing

9 questions Answer LLM: offline-sim-v1 ai_context_benchmark_v1

Of 9 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 0 · AIPM 5 · multi-layer 1 · unanswered 2. Descriptive counts only; no overall ranking.

Of 9 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 0 · AIPM 5 · multi-layer 1 · unanswered 2. Descriptive counts only; no overall ranking. Each question shows: sufficient layer(s) → retrieved slice → measured input tokens → evidence. Readers interpret; the report does not rank formats. Semantic Redundancy 100% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

Gold mode · Mixed independent + AIPM consistency

Most Understanding/Retrieval gold comes from HTML. Some questions remain AIPM self-consistency checks (tagged) and are excluded from Understanding when possible. Matching a machine card against its own fields is not evidence that that layer outperforms HTML.

Execution provenance

Methodology
AI Context Benchmark Methodology v1.0
Planner Spec
v1.0
Execution Protocol
v1.0
Question pack
universal_v1
Score version
aipm_benchmark_score_v10
Engine
aipm_benchmark_v4
Answer LLM
simulation / offline-sim-v1 (single provider this lab)
Model snapshot
2026-07
Run id
79d55372-7ea9-4458-ae20-85831faa742a

Question outcomes

Per question: which layers were sufficient, and what was the lowest measured retrieval cost among them. No overall winner.

1

Lowest cost: HTML

0

Lowest cost: Schema

5

Lowest cost: AIPM

1

Multi-layer

2

Unanswered

Manifest design

Semantic Redundancy · 100%

Semantic Redundancy 100% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

  • Merge or differentiate `purpose` and `abstract` (100% overlap).
Field A Field B Overlap
purpose abstract 100%

Structural size (secondary)

Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.

HTML · 5,802 chars
Schema · 14,923 chars
AIPM · 8,641 chars

Routing helper (secondary)

Illustrative card-first vs HTML-always simulation — prefer per-question measured cost above. Not a ranking.

Card-first retrieval would use ~35% fewer tokens than HTML-always on this pack (7 answered from machine card, 2 escalated to HTML). Descriptive only.

Orientation pack

7 questions · all layers scored

  • HTML 2/7
  • Schema 5/7
  • AIPM 5/7

Depth pack

0 questions · all layers scored

  • HTML 0/0
  • Schema 0/0
  • AIPM 0/0

Full-context pack metrics (secondary)

These measure the whole file fed to the model this run — not the minimum slice needed per question.

HTML

3/9 matched

2,700 full-pack tokens

Schema

6/9 matched

3,640 full-pack tokens

AIPM

7/9 matched

1,472 full-pack tokens

Full-pack resource table (secondary)

Whole-file context fed this run. Prefer Minimal Retrieval Cost on each question.

Metric HTML Schema AIPM
Coverage (matched) 3/9 6/9 7/9
Context size 5,802 chars 14,923 chars 8,641 chars
Total tokens 2,700 3,640 1,472
Tokens / correct answer 900 607 210
Est. cost / correct answer $0.000153 $0.000103 $0.000042
Matched per 1k tokens 1.111 1.648 4.755
Median latency 994 ms 784 ms 554 ms
Est. cost (USD) $0.00046 $0.00062 $0.00029

Six independent scores

Answer Efficiency is the primary cost lens. Accuracy axes remain for research — no combined total or winner.

Answer Efficiency

Matched answers per 1k tokens (and cost per match). The primary efficiency axis — not raw accuracy.

  • HTML 23.4
  • Schema 34.7
  • AIPM 100

Understanding

Can this layer convey what the page is about — using independent HTML gold?

  • HTML 40
  • Schema 80
  • AIPM 80

Retrieval

Can this layer surface shared facts (location, contact, hours, pricing, FAQ, CTA)?

  • HTML
  • Schema
  • AIPM

Evidence

Answer quality vs independent gold (score strength). Partial credit counts; UNKNOWN scores zero unless gold is UNKNOWN.

  • HTML 31.3
  • Schema 67.8
  • AIPM 78.9

Metadata

Language, page kind, and freshness from page signals.

  • HTML 0
  • Schema 50
  • AIPM 50

Compression

Information delivered per token and context size. Higher means more matched answers for less context cost.

  • HTML 23.4
  • Schema 34.7
  • AIPM 100

Coverage map

AIPM matched 7 question(s); HTML matched 3. HTML/Schema (or a gap) still needed on Q1, Q3. No layer matched gold on Q1, Q3. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 3/9 matched (≈900 tok/match). Schema 6/9 matched (≈607 tok/match). AIPM 7/9 matched (≈210 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

AIPM matched

Q2, Q4, Q5, Q6, Q7, Q8, Q9

Alone: —

AIPM insufficient

Q1, Q3

HTML/Schema needed or all layers missed

Unanswered by all

Q1, Q3

HTML layer

Visible page text after stripping AIPM sidecars and JSON-LD. Measures what prose alone can answer.

Context fed: 5,802 chars

Tokens: 2,700 · median 994 ms

Matched this pack: 3/9

Stronger on

Weaker on

Metadata (0) · Evidence (31.3) · Compression (23.4) · Answer Efficiency (23.4)

Schema layer

JSON-LD structured data with minimal page chrome. Measures what schema markup can answer.

Context fed: 14,923 chars

Tokens: 3,640 · median 784 ms

Matched this pack: 6/9

Stronger on

Understanding (80)

Weaker on

Compression (34.7) · Answer Efficiency (34.7)

AIPM layer

AI Page Manifest (.ai.json) only. Measures what the machine layer can answer without HTML.

Context fed: 8,641 chars

Tokens: 1,472 · median 554 ms

Matched this pack: 7/9

Stronger on

Understanding (80) · Evidence (78.9) · Compression (100) · Answer Efficiency (100)

Weaker on

When to use which layer

AIPM complements HTML — it does not replace full-page prose.

Scenario Recommended Why
Fast orientation (title, purpose, brand, intent) Compare machine card → HTML fallback on this run Machine-card pack: 1,472 tok · $0.00029. Routing sim saved ~35% tokens vs HTML-always.
Deep content / research (prose facts, process detail) HTML (with optional machine orientation) HTML pack: 2,700 tok · $0.00046. Use when depth needs body prose.
Structured entity pulls (org, location, typed fields) Schema.org Schema pack: 3,640 tok · $0.00062. Dense JSON-LD tends to score well here.

Findings

  • Of 9 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 0 · AIPM 5 · multi-layer 1 · unanswered 2. Descriptive counts only; no overall ranking.
  • Semantic Redundancy 100% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.
  • Manifest design: Merge or differentiate `purpose` and `abstract` (100% overlap).
  • Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.
  • AIPM matched 7 question(s); HTML matched 3. HTML/Schema (or a gap) still needed on Q1, Q3. No layer matched gold on Q1, Q3. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 3/9 matched (≈900 tok/match). Schema 6/9 matched (≈607 tok/match). AIPM 7/9 matched (≈210 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

Question-by-question layer analysis

Which layer(s) could answer; which need more or different context; minimum context fed this run.

Q1 · Understanding · orientation

What is the primary topic of this page?

Gold: Pricing built for every stage

Gold source: primaryTopic · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.h1|title est no
Schema No typed Schema field for this question. est no
AIPM primaryTopic est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

insufficient

Not clearly stated in the provided context.

score 0 · 5,802 chars context · 258 in-tokens

HTML did not answer from 5802 chars of context — additional or different layer context needed.

Schema

insufficient

AI visibility pricing

score 29 · 14,923 chars context · 348 in-tokens

Schema did not answer from 14923 chars of context — additional or different layer context needed.

AIPM

insufficient

AI visibility pricing

score 29 · 8,641 chars context · 131 in-tokens

AIPM did not answer from 8641 chars of context — additional or different layer context needed.

Q2 · Retrieval · pricing

Name one pricing-related fact from this page.

Gold: Plans range from Free to Enterprise.

Gold source: keyFacts[0] · unknown

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No typed Schema field for this question. est yes
AIPM No minimal AIPM slice identified for this question. est yes

HTML

insufficient

Not clearly stated in the provided context.

score 0 · 5,802 chars context · 258 in-tokens

HTML did not answer from 5802 chars of context — additional or different layer context needed.

Schema

sufficient

Plans range from Free to Enterprise.

score 100 · 14,923 chars context · 348 in-tokens

Schema answered using 14923 chars of layer context (minimum fed this run).

AIPM

sufficient

Plans range from Free to Enterprise.

score 100 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Q3 · Metadata · orientation

What is the content intent?

Gold: informational

Gold source: contentIntent · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_signals est no
Schema No typed Schema field for this question. est no
AIPM contentIntent est no

Root cause · Wrong contentIntent

Set AIPM contentIntent to match page intent (commercial vs informational).

HTML

insufficient

commercial

score 0 · 5,802 chars context · 258 in-tokens

HTML did not answer from 5802 chars of context — additional or different layer context needed.

Schema

insufficient

Not clearly stated in the provided context.

score 0 · 14,923 chars context · 348 in-tokens

Schema did not answer from 14923 chars of context — additional or different layer context needed.

AIPM

insufficient

commercial

score 0 · 8,641 chars context · 131 in-tokens

AIPM did not answer from 8641 chars of context — additional or different layer context needed.

Q4 · Metadata · orientation

What language is this page in?

Gold: en

Gold source: inLanguage · html_independent

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost · lowest cost AIPM · 1 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lang est no
Schema jsonld.inLanguage 1 est yes
AIPM inLanguage 1 est yes

HTML

insufficient

Not clearly stated in the provided context.

score 0 · 5,802 chars context · 258 in-tokens

HTML did not answer from 5802 chars of context — additional or different layer context needed.

Schema

sufficient

en

score 100 · 14,923 chars context · 348 in-tokens

Schema answered using 14923 chars of layer context (minimum fed this run).

AIPM

sufficient

en

score 100 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Q5 · Retrieval · local_service

Which location or city is mentioned in the key facts?

Gold: Plans

Gold source: keyFacts · unknown

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost HTML · 2 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.address|prose 2 est yes
Schema jsonld.block est no
AIPM Field missing in AIPM — cannot answer from this layer alone. est yes

HTML

sufficient

Plans

score 100 · 5,802 chars context · 258 in-tokens

HTML answered using 5802 chars of layer context (minimum fed this run).

Schema

insufficient

Not clearly stated in the provided context.

score 0 · 14,923 chars context · 348 in-tokens

Schema did not answer from 14923 chars of context — additional or different layer context needed.

AIPM

sufficient

Plans

score 100 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Q6 · Understanding · orientation

What is the purpose of this page?

Gold: Subscription plans from Free through Enterprise for AI visibility monitoring, GEO optimization, and discovery file deployment.

Gold source: purpose · html_independent

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost · lowest cost AIPM · 32 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.meta_description est no
Schema jsonld.description 32 est yes
AIPM purpose 32 est yes

HTML

insufficient

UNKNOWN

score 0 · 5,802 chars context · 258 in-tokens

HTML did not answer from 5802 chars of context — additional or different layer context needed.

Schema

sufficient

Subscription plans from Free through Enterprise for AI visibility monitoring, GEO optimization, and discovery file deployment.

score 100 · 14,923 chars context · 348 in-tokens

Schema answered using 14923 chars of layer context (minimum fed this run).

AIPM

sufficient

Subscription plans from Free through Enterprise for AI visibility monitoring, GEO optimization, and discovery file deployment.

score 100 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Q7 · Understanding · orientation

What is the page title?

Gold: Pricing — AigeoRadar

Gold source: title · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost AIPM · 2 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title 5 est yes
Schema jsonld.name 3 est yes
AIPM title 2 est yes

HTML

sufficient

Pricing

score 82 · 5,802 chars context · 258 in-tokens

HTML answered using 5802 chars of layer context (minimum fed this run).

Schema

sufficient

Pricing

score 82 · 14,923 chars context · 348 in-tokens

Schema answered using 14923 chars of layer context (minimum fed this run).

AIPM

sufficient

Pricing

score 82 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Q8 · Understanding · orientation

Who is the publisher or brand?

Gold: AigeoRadar

Gold source: publisher.name · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost AIPM · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title_brand 5 est yes
Schema jsonld.name 3 est yes
AIPM publisher 3 est yes

HTML

sufficient

AigeoRadar

score 100 · 5,802 chars context · 258 in-tokens

HTML answered using 5802 chars of layer context (minimum fed this run).

Schema

sufficient

AigeoRadar

score 100 · 14,923 chars context · 348 in-tokens

Schema answered using 14923 chars of layer context (minimum fed this run).

AIPM

sufficient

AigeoRadar

score 100 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Q9 · Understanding · orientation

Summarize the page in one sentence.

Gold: Subscription plans from Free through Enterprise for AI visibility monitoring, GEO optimization, and discovery file deployment.

Gold source: abstract · html_independent

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost · lowest cost AIPM · 32 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.meta_description est no
Schema jsonld.description 32 est yes
AIPM abstract 32 est yes

HTML

insufficient

Not clearly stated in the provided context.

score 0 · 5,802 chars context · 258 in-tokens

HTML did not answer from 5802 chars of context — additional or different layer context needed.

Schema

sufficient

Subscription plans from Free through Enterprise for AI visibility monitoring, GEO optimization, and discovery file deployment.

score 100 · 14,923 chars context · 348 in-tokens

Schema answered using 14923 chars of layer context (minimum fed this run).

AIPM

sufficient

Subscription plans from Free through Enterprise for AI visibility monitoring, GEO optimization, and discovery file deployment.

score 100 · 8,641 chars context · 131 in-tokens

AIPM answered using 8641 chars of layer context (minimum fed this run).

Methodology

  • AI Context Benchmark: Planner → Slice → LLM with measured API input tokens.
  • Reports describe sufficient layers, slice size, cost, and evidence — they do not declare a winning format.
  • Layers under test today: HTML, Schema.org, AIPM (extensible to RSS, Markdown, PDF, …).
  • Engine aipm_benchmark_v1 · Mon, Jul 27, 2026 11:47 PM · aipm_benchmark_score_v10

Share

Per-question context chain across layers.