AigeoRadar

AigeoRadar · AI Context Benchmark

AI Context Benchmark

Question → sufficient layer(s) → retrieved slice → measured cost → evidence

Ankara Villa Tadilat ve Anahtar Teslim Villa Yenileme

https://ankaraanahtarteslim.com.tr/ankara-villa-tadilat-ve-anahtar-teslim-villa-yenileme-firmasi/

13 questions Answer LLM: gpt-4o-mini ai_context_benchmark_v1

Of 13 questions: lowest measured retrieval cost among sufficient layers — HTML 3 · Schema 0 · AIPM 5 · multi-layer 0 · unanswered 5. Descriptive counts only; no overall ranking.

Of 13 questions: lowest measured retrieval cost among sufficient layers — HTML 3 · Schema 0 · AIPM 5 · multi-layer 0 · unanswered 5. Descriptive counts only; no overall ranking. Each question shows: sufficient layer(s) → retrieved slice → measured input tokens → evidence. Readers interpret; the report does not rank formats. Semantic Redundancy 59% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

Gold mode · Mixed independent + AIPM consistency

Most Understanding/Retrieval gold comes from HTML. Some questions remain AIPM self-consistency checks (tagged) and are excluded from Understanding when possible. Matching a machine card against its own fields is not evidence that that layer outperforms HTML.

Execution provenance

Methodology
AI Context Benchmark Methodology v1.0
Planner Spec
v1.0
Execution Protocol
v1.0
Question pack
universal_v1
Score version
aipm_benchmark_score_v10
Engine
aipm_benchmark_v4
Answer LLM
openai / gpt-4o-mini (single provider this lab)
Model snapshot
2026-07
Run id
2744acfb-f5cd-4650-80c2-d24ab3717c83

Question outcomes

Per question: which layers were sufficient, and what was the lowest measured retrieval cost among them. No overall winner.

3

Lowest cost: HTML

0

Lowest cost: Schema

5

Lowest cost: AIPM

0

Multi-layer

5

Unanswered

Manifest design

Semantic Redundancy · 59%

Semantic Redundancy 59% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

  • Merge or differentiate `purpose` and `abstract` (100% overlap).
Field A Field B Overlap
purpose abstract 100%
purpose keyFacts 47%
abstract keyFacts 47%
abstract title 42%

Structural size (secondary)

Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.

HTML · 12,006 chars
Schema · 225 chars
AIPM · 10,574 chars

Routing helper (secondary)

Illustrative card-first vs HTML-always simulation — prefer per-question measured cost above. Not a ranking.

Card-first retrieval would use ~8% fewer tokens than HTML-always on this pack (6 answered from machine card, 7 escalated to HTML). Descriptive only.

Orientation pack

9 questions · all layers scored

  • HTML 3/9
  • Schema 2/9
  • AIPM 5/9

Depth pack

4 questions · all layers scored

  • HTML 2/4
  • Schema 1/4
  • AIPM 1/4

Full-context pack metrics (secondary)

These measure the whole file fed to the model this run — not the minimum slice needed per question.

HTML

5/13 matched

50,456 full-pack tokens

Schema

3/13 matched

2,137 full-pack tokens

AIPM

6/13 matched

41,262 full-pack tokens

Full-pack resource table (secondary)

Whole-file context fed this run. Prefer Minimal Retrieval Cost on each question.

Metric HTML Schema AIPM
Coverage (matched) 5/13 3/13 6/13
Context size 12,006 chars 225 chars 10,574 chars
Total tokens 50,456 2,137 41,262
Tokens / correct answer 10,091 712 6,877
Est. cost / correct answer $0.001531 $0.000123 $0.001048
Matched per 1k tokens 0.099 1.404 0.145
Median latency 1,186 ms 878 ms 1,126 ms
Est. cost (USD) $0.00766 $0.00037 $0.00629

Six independent scores

Answer Efficiency is the primary cost lens. Accuracy axes remain for research — no combined total or winner.

Answer Efficiency

Matched answers per 1k tokens (and cost per match). The primary efficiency axis — not raw accuracy.

  • HTML 7.1
  • Schema 100
  • AIPM 10.3

Understanding

Can this layer convey what the page is about — using independent HTML gold?

  • HTML 20
  • Schema 20
  • AIPM 60

Retrieval

Can this layer surface shared facts (location, contact, hours, pricing, FAQ, CTA)?

  • HTML 50
  • Schema 25
  • AIPM 25

Evidence

Answer quality vs independent gold (score strength). Partial credit counts; UNKNOWN scores zero unless gold is UNKNOWN.

  • HTML 46.6
  • Schema 30.2
  • AIPM 46.3

Metadata

Language, page kind, and freshness from page signals.

  • HTML 66.7
  • Schema 33.3
  • AIPM 66.7

Compression

Information delivered per token and context size. Higher means more matched answers for less context cost.

  • HTML 7.1
  • Schema 100
  • AIPM 10.3

Coverage map

AIPM matched 6 question(s) (alone on Q3, Q5, Q8); HTML matched 5. HTML/Schema (or a gap) still needed on Q2, Q4, Q7, Q9, Q11, Q12, Q13. No layer matched gold on Q2, Q4, Q9, Q11, Q12. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 5/13 matched (≈10,091 tok/match). Schema 3/13 matched (≈712 tok/match). AIPM 6/13 matched (≈6,877 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

AIPM matched

Q1, Q3, Q5, Q6, Q8, Q10

Alone: Q3, Q5, Q8

AIPM insufficient

Q2, Q4, Q7, Q9, Q11, Q12, Q13

HTML/Schema needed or all layers missed

Unanswered by all

Q2, Q4, Q9, Q11, Q12

HTML layer

Visible page text after stripping AIPM sidecars and JSON-LD. Measures what prose alone can answer.

Context fed: 12,006 chars

Tokens: 50,456 · median 1,186 ms

Matched this pack: 5/13

Stronger on

Weaker on

Understanding (20) · Compression (7.1) · Answer Efficiency (7.1)

Schema layer

JSON-LD structured data with minimal page chrome. Measures what schema markup can answer.

Context fed: 225 chars

Tokens: 2,137 · median 878 ms

Matched this pack: 3/13

Stronger on

Compression (100) · Answer Efficiency (100)

Weaker on

Understanding (20) · Retrieval (25) · Metadata (33.3) · Evidence (30.2)

AIPM layer

AI Page Manifest (.ai.json) only. Measures what the machine layer can answer without HTML.

Context fed: 10,574 chars

Tokens: 41,262 · median 1,126 ms

Matched this pack: 6/13

Stronger on

Weaker on

Retrieval (25) · Compression (10.3) · Answer Efficiency (10.3)

When to use which layer

AIPM complements HTML — it does not replace full-page prose.

Scenario Recommended Why
Fast orientation (title, purpose, brand, intent) Compare machine card → HTML fallback on this run Machine-card pack: 41,262 tok · $0.00629. Routing sim saved ~8% tokens vs HTML-always.
Deep content / research (prose facts, process detail) HTML (with optional machine orientation) HTML pack: 50,456 tok · $0.00766. Use when depth needs body prose.
Structured entity pulls (org, location, typed fields) Schema.org Schema pack: 2,137 tok · $0.00037. Dense JSON-LD tends to score well here.

Findings

  • Of 13 questions: lowest measured retrieval cost among sufficient layers — HTML 3 · Schema 0 · AIPM 5 · multi-layer 0 · unanswered 5. Descriptive counts only; no overall ranking.
  • Semantic Redundancy 59% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.
  • Manifest design: Merge or differentiate `purpose` and `abstract` (100% overlap).
  • Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.
  • AIPM matched 6 question(s) (alone on Q3, Q5, Q8); HTML matched 5. HTML/Schema (or a gap) still needed on Q2, Q4, Q7, Q9, Q11, Q12, Q13. No layer matched gold on Q2, Q4, Q9, Q11, Q12. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 5/13 matched (≈10,091 tok/match). Schema 3/13 matched (≈712 tok/match). AIPM 6/13 matched (≈6,877 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

Question-by-question layer analysis

Which layer(s) could answer; which need more or different context; minimum context fed this run.

Q1 · Understanding · orientation

What is the page title?

Gold: Ankara Villa Tadilat ve Yenileme | Nokta Yapı – Anahtar Teslim Hizmet – Ankara Anahtar Teslim

Gold source: html.title · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost AIPM · 18 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title 24 est yes
Schema No JSON-LD on page. est yes
AIPM title 18 est yes

HTML

sufficient

Ankara Villa Tadilat ve Yenileme | Nokta Yapı – Anahtar Teslim Hizmet – Ankara Anahtar Teslim

score 100 · 12,006 chars context · 3,859 in-tokens

HTML answered using 12006 chars of layer context (minimum fed this run).

Schema

sufficient

The page title is "Ankara Villa Tadilat ve Yenileme | Nokta Yapı – Anahtar Teslim Hizmet – Ankara Anahtar Teslim."

score 95 · 225 chars context · 149 in-tokens

Schema answered using 225 chars of layer context (minimum fed this run).

AIPM

sufficient

The page title is "Ankara Villa Tadilat ve Yenileme | Nokta Yapı – Anahtar Teslim Hizmet."

score 95 · 10,574 chars context · 3,150 in-tokens

AIPM answered using 10574 chars of layer context (minimum fed this run).

Q2 · Understanding · orientation

What is the primary topic of this page?

Gold: Ankara Villa Tadilat ve Yenileme

Gold source: html.h1 · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.h1|title est no
Schema No JSON-LD on page. est no
AIPM primaryTopic est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

insufficient

The primary topic of this page is Ankara villa renovation and remodeling services offered by Nokta Yapı.

score 31 · 12,006 chars context · 3,862 in-tokens

HTML did not answer from 12006 chars of context — additional or different layer context needed.

Schema

insufficient

The primary topic of this page is villa renovation and renewal services in Ankara.

score 36 · 225 chars context · 152 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

insufficient

The primary topic of this page is "Ankara Anahtar Teslim."

score 22 · 10,574 chars context · 3,153 in-tokens

AIPM did not answer from 10574 chars of context — additional or different layer context needed.

Q3 · Understanding · orientation

What is the purpose of this page?

Gold: Ankara Anahtar Teslim: Ankara Villa Tadilat ve Yenileme — page overview from HTML signals.

Gold source: html.page_summary · html_independent

Sufficient: AIPM · Needs more / other context: HTML, Schema

Estimated retrieval cost · lowest cost AIPM · 40 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema No JSON-LD on page. est no
AIPM purpose 40 est yes

HTML

insufficient

The purpose of this page is to provide information about villa renovation and remodeling services offered by Nokta Yapı in Ankara.

score 25 · 12,006 chars context · 3,861 in-tokens

HTML did not answer from 12006 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 225 chars context · 151 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

sufficient

The purpose of the page is to provide information about "Ankara Villa Tadilat ve Anahtar Teslim Villa Yenileme Firması | Nokta Yapı" in various neighborhoods of Ankara.

score 63 · 10,574 chars context · 3,152 in-tokens

AIPM answered using 10574 chars of layer context (minimum fed this run).

Q4 · Understanding · orientation

Summarize the page in one sentence.

Gold: Ankara Anahtar Teslim: Ankara Villa Tadilat ve Yenileme — page overview from HTML signals.

Gold source: html.page_summary · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema No JSON-LD on page. est no
AIPM abstract est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

insufficient

Nokta Yapı offers professional villa renovation and turnkey services in Ankara, providing customized solutions with a focus on quality and customer satisfaction.

score 25 · 12,006 chars context · 3,862 in-tokens

HTML did not answer from 12006 chars of context — additional or different layer context needed.

Schema

insufficient

The page is about renovation and renewal services for villas in Ankara offered by Nokta Yapı, emphasizing their turnkey service.

score 13 · 225 chars context · 152 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

insufficient

The page provides information about Nokta Yapı, a company offering villa renovation and turnkey services in various neighborhoods of Ankara, including Çayyolu, Yaşamkent, and Gölbaşı.

score 25 · 10,574 chars context · 3,153 in-tokens

AIPM did not answer from 10574 chars of context — additional or different layer context needed.

Q5 · Understanding · orientation

Who is the publisher or brand?

Gold: Ankara Anahtar Teslim

Gold source: html.jsonld|title · html_independent

Sufficient: AIPM · Needs more / other context: HTML, Schema

Estimated retrieval cost · lowest cost AIPM · 6 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title_brand est no
Schema No JSON-LD on page. est no
AIPM publisher 6 est yes

HTML

insufficient

Öz Balıkçılar İnşaat.

score 0 · 12,006 chars context · 3,860 in-tokens

HTML did not answer from 12006 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 225 chars context · 150 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

sufficient

The publisher is Ankara Anahtar Teslim.

score 88 · 10,574 chars context · 3,151 in-tokens

AIPM answered using 10574 chars of layer context (minimum fed this run).

Q6 · Metadata · orientation

What language is this page in?

Gold: tr

Gold source: html.lang|script · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 1 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lang 2 est yes
Schema No JSON-LD on page. est no
AIPM inLanguage 1 est yes

HTML

sufficient

The page is in Turkish.

score 88 · 12,006 chars context · 3,860 in-tokens

HTML answered using 12006 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 225 chars context · 150 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

sufficient

The page is in Turkish (tr).

score 88 · 10,574 chars context · 3,151 in-tokens

AIPM answered using 10574 chars of layer context (minimum fed this run).

Q7 · Metadata · orientation

What is the content intent?

Gold: commercial

Gold source: html.heuristic · html_independent

Sufficient: HTML, Schema · Partial: AIPM

Estimated retrieval cost · lowest cost HTML · 100 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_signals 100 est yes
Schema No JSON-LD on page. est yes
AIPM contentIntent est no

Root cause · Wrong contentIntent

Set AIPM contentIntent to match page intent (commercial vs informational).

HTML

sufficient

The content intent is to promote and provide information about villa renovation and remodeling services in Ankara by Nokta Yapı.

score 88 · 12,006 chars context · 3,859 in-tokens

HTML answered using 12006 chars of layer context (minimum fed this run).

Schema

sufficient

The content intent is to provide information about renovation and renewal services in Ankara.

score 88 · 225 chars context · 149 in-tokens

Schema answered using 225 chars of layer context (minimum fed this run).

AIPM

partial

The content intent is informational.

score 40 · 10,574 chars context · 3,150 in-tokens

AIPM was only partially aligned with gold from 10574 chars of context.

Q8 · Metadata · orientation

What page type is declared in the machine layer?

Gold: blog

Gold source: pageType · aipm_consistency

Sufficient: AIPM

Estimated retrieval cost · lowest cost AIPM · 1 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No JSON-LD on page. est no
AIPM pageType 1 est yes

HTML

skipped

UNKNOWN

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

Schema

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

AIPM

sufficient

The page type declared in the machine layer is "blog."

score 88 · 10,574 chars context · 3,154 in-tokens

AIPM answered using 10574 chars of layer context (minimum fed this run).

Q9 · Understanding · orientation

Who is the intended audience declared in AIPM?

Gold: Ev sahipleri, daire sahipleri, villa sahipleri, iş yeri sahipleri, ofis yöneticileri, mağaza işletmeleri, eczane sahipleri, diş klinikleri, yatırım amaçlı gayrimenkul sahipleri, müteahhitler, mimarlar, iç mimarlar, site yönetimleri, Ankara'da profesyonel anahtar teslim tadilat

Gold source: audience · aipm_consistency

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No JSON-LD on page. est no
AIPM audience est no

Root cause · Self-consistency miss

AIPM failed to read its own declared field — check serialization.

HTML

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

Schema

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

AIPM

insufficient

The intended audience includes homeowners, apartment owners, villa owners, business owners, office managers, store operators, pharmacy owners, dental clinics, real estate investors, contractors, architects, interior designers, site management, and professional renovation services in Ankara.

score 11 · 10,574 chars context · 3,155 in-tokens

AIPM did not answer from 10574 chars of context — additional or different layer context needed.

Q10 · Retrieval · depth

Which location or city is mentioned for this business?

Gold: Ankara

Gold source: html.address|prose · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 2 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.address|prose 2 est yes
Schema No JSON-LD on page. est yes
AIPM Field missing in AIPM — cannot answer from this layer alone. est yes

HTML

sufficient

Ankara.

score 100 · 12,006 chars context · 3,863 in-tokens

HTML answered using 12006 chars of layer context (minimum fed this run).

Schema

sufficient

Ankara.

score 100 · 225 chars context · 153 in-tokens

Schema answered using 225 chars of layer context (minimum fed this run).

AIPM

sufficient

Ankara is mentioned for this business.

score 88 · 10,574 chars context · 3,154 in-tokens

AIPM answered using 10574 chars of layer context (minimum fed this run).

Q11 · Retrieval · depth

According to the page, what does it say about: "Ev, ofis ve iş yerlerinizde profesyonel ve güven…"?

Gold: Ev, ofis ve iş yerlerinizde profesyonel ve güvenilir çözümler için hemen bize ulaşın.

Gold source: html.prose · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_span est no
Schema No JSON-LD on page. est no
AIPM Prose fact not present as a discrete AIPM field — escalate to HTML. est no

Root cause · Depth gap (expected)

Prose/detail questions belong to HTML — escalate after AIPM orientation.

HTML

insufficient

UNKNOWN.

score 0 · 12,006 chars context · 3,882 in-tokens

HTML did not answer from 12006 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 225 chars context · 172 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

insufficient

UNKNOWN.

score 0 · 10,574 chars context · 3,173 in-tokens

AIPM did not answer from 10574 chars of context — additional or different layer context needed.

Q12 · Retrieval · depth

According to the page, what does it say about: "10 yılı aşkın tecrübemiz ve uzman ekibimizle vil…"?

Gold: 10 yılı aşkın tecrübemiz ve uzman ekibimizle villanızı baştan aşağı yeniliyor, hayalinizdeki yaşam alanını teslim ediyoruz.

Gold source: html.prose · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_span est no
Schema No JSON-LD on page. est no
AIPM Prose fact not present as a discrete AIPM field — escalate to HTML. est no

Root cause · Depth gap (expected)

Prose/detail questions belong to HTML — escalate after AIPM orientation.

HTML

insufficient

It states that they have over 10 years of experience and a specialized team to renovate villas.

score 7 · 12,006 chars context · 3,882 in-tokens

HTML did not answer from 12006 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 225 chars context · 172 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

insufficient

UNKNOWN.

score 0 · 10,574 chars context · 3,173 in-tokens

AIPM did not answer from 10574 chars of context — additional or different layer context needed.

Q13 · Retrieval · depth

According to the page, what does it say about: "Ücretsiz Keşif ve Proje: Uzman ekibimiz villanız…"?

Gold: Ücretsiz Keşif ve Proje: Uzman ekibimiz villanıza gelir, ihtiyaçlarınızı dinler ve ücretsiz keşif yaparak size özel tadilat planı hazırlar.

Gold source: html.prose · html_independent

Sufficient: HTML · Needs more / other context: Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 35 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_span 35 est yes
Schema No JSON-LD on page. est no
AIPM Prose fact not present as a discrete AIPM field — escalate to HTML. est no

Root cause · Depth gap (expected)

Prose/detail questions belong to HTML — escalate after AIPM orientation.

HTML

sufficient

Ücretsiz keşif ve proje, uzman ekibin villaya gelerek ihtiyaçları dinleyip özel tadilat planı hazırlamasını içerir.

score 48 · 12,006 chars context · 3,883 in-tokens

HTML answered using 12006 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 225 chars context · 173 in-tokens

Schema did not answer from 225 chars of context — additional or different layer context needed.

AIPM

insufficient

UNKNOWN.

score 0 · 10,574 chars context · 3,174 in-tokens

AIPM did not answer from 10574 chars of context — additional or different layer context needed.

Methodology

  • AI Context Benchmark: Planner → Slice → LLM with measured API input tokens.
  • Reports describe sufficient layers, slice size, cost, and evidence — they do not declare a winning format.
  • Layers under test today: HTML, Schema.org, AIPM (extensible to RSS, Markdown, PDF, …).
  • Engine aipm_benchmark_v4 · Wed, Jul 29, 2026 12:37 PM · aipm_benchmark_score_v10

Share

Per-question context chain across layers.