---
title: "Build an Entity Gap Audit With Schema.org and Google's NLP API - AstroDev"
description: "An entity gap audit measures the distance between the entities you declare in schema and the ones Google's NLP actually recognizes. Here is the pipeline."
url: "https://astrodev.carlosarias.com/blog/guides/entity-gap-audit-schema-nlp"
---

[Guides](/categories/guides)

# Build an Entity Gap Audit With Schema.org and Google's NLP API

An entity gap audit measures the distance between the entities you declare in schema and the ones Google's NLP actually recognizes. Here is the pipeline.

  [Carlos Arias](/authors/carlos-arias) · September 10, 2026  · 7 min read

![Declared schema entities joined against the entities Google's NLP recognizes, diffed by MID.](/_astro/cover.VX9J5NUt_ZhcvWz.webp)

*Declared schema entities joined against the entities Google's NLP recognizes, diffed by MID. AI-generated illustration by Carlos Arias .*

      On-brand editorial cover for an article titled "Build an Entity Gap Audit With Schema.org and Google's NLP API". Sophisticated, minimal conceptual illustration on a very dark ink background (#111318) with a single restrained warm accent glow. High-end business-publication aesthetic, subtle depth, cinematic soft light. No text, no words, no letters, no logos, no UI labels.  Prompt sent to Higgsfield · nano_banana_pro · 3:2

An entity gap audit measures one distance: the entities you declare in your structured data versus the entities Google’s language models actually recognize on the same page. Those two sets are rarely equal. You mark up an Organization, a Product, an author, a place, and you assume the machine reads them the way you wrote them. It often does not. The audit turns that assumption into a query you can run, version, and rerun after every content change.

This is a scriptable pipeline, not a dashboard. Schema markup goes in, the Google Cloud Natural Language API extracts what it sees, and an agentic coding tool joins the two into a small knowledge graph you can diff. The idea traces to a September 2026 SMX session on entity gaps in content strategy (Search Engine Land, 2026). What follows is that pitch built out as an actual pipeline, with the failure modes an engineer hits along the way.

## What an entity gap audit actually measures

Start with the vocabulary, because the word “entity” gets used three different ways in one breath. There is the entity you declare: a JSON-LD node with an @type, a name, and ideally a sameAs. There is the entity a language model extracts from your rendered text. And there is the entity in Google’s production Knowledge Graph, the one with a stable identifier that powers panels and answers.

The gap that matters sits between the first two, read as a proxy for the third. You cannot query Google’s ranking-time graph directly. What you can do is take a documented model, the Cloud Natural Language API, and treat its output as a reasonable stand-in for “what a Google-scale extractor understands from this page.” That framing is the whole audit. It is also its main caveat, and I will come back to it.

An entity gap shows up in two directions. Entities you declare but the model never surfaces suggest your markup is decorative, disconnected from the prose. Entities the model ranks as central but you never declared suggest a page that is about something your structured data ignores. Both are fixable. Neither is visible without measuring.

## Step one: extract the entities you declare

Pull the JSON-LD from the rendered page, not the raw template. Parse each node and keep three fields per entity: the @type, the name, and every URL in sameAs. The sameAs property is the load-bearing one. It links your entity to an authoritative reference such as a Wikipedia or Wikidata URL, which is how a search engine disambiguates your “Mercury” from the planet and the element (schema.org).

Normalize the result into a flat list of declared entities with their external references. This is your ground truth for the audit: the set of things you have told the machine your page is about. Keep it in the same store you will use for the extracted set, because the join between them is the entire point.

One trap here. If your schema is injected client-side by a tag manager or a framework hydration step, a fetch of the HTML source will miss it. Render the page in a headless browser first, then read the DOM. Otherwise your declared set comes back empty and every entity reads as a gap.

## Step two: extract the entities Google’s NLP recognizes

Send the page’s visible text to the Natural Language API’s analyzeEntities method. For each entity it returns a type, a salience score between 0 and 1 that estimates how central the entity is to the document, and a metadata block. When the API can resolve an entity to Google’s Knowledge Graph, that metadata carries a mid, the Machine ID, and often a wikipedia_url (Cloud Natural Language, Entity reference).

That MID is the join key you have been waiting for. It is a stable, language-independent handle for a Knowledge Graph node, so two entities with the same MID are the same thing even if the surface strings differ. Where a MID is present, you can match a declared entity to a recognized one with confidence. Where it is absent, you fall back to fuzzy name matching and accept the noise.

Cost is real but small at audit scale. Entity analysis is free for the first 5,000 units a month, then $1.00 per 1,000 units, where a unit is up to 1,000 characters, as listed on the Cloud Natural Language pricing page as of September 2026. A 2,000-word article is roughly a dozen units. You can audit a few hundred pages for the price of a coffee, which is the point of scripting it rather than buying a seat in a tool.

## Step three: join the two sets into a queryable graph

Now you have two tables. Declared entities keyed by name and sameAs, and recognized entities keyed by MID, name, and salience. Load both into something you can query. SQLite is enough. A property graph is nicer if you want to walk relationships later, but do not reach for it on day one.

Model the join as a left-and-right comparison rather than a single merge. For every declared entity, ask whether a recognized entity shares its MID or resolves to the same reference URL. For every recognized entity above a salience threshold, ask whether you declared it. The rows that fail either test are your audit output. This is where a schema.org knowledge graph stops being a marketing phrase and becomes a table you can SELECT against.

Set the salience threshold deliberately. Salience is document-relative, so it tells you what this page is mostly about, not how authoritative the entity is on the web. A threshold around 0.02 to 0.05 filters out the passing mentions without discarding secondary topics. Tune it against pages whose subject you already know, then keep it fixed so audits stay comparable over time. That discipline of holding one variable steady is the same one that separates a real test from a story told afterward, which I covered in technical SEO testing.

## Step four: add the competitor diff

The audit gets sharper when you run it against pages that already rank for your target. Extract the recognized entities from two or three competitor URLs the same way, and diff their high-salience set against yours. Entities that show up as central across competitors but are missing from your page are the topical coverage gap. This is the entity SEO audit move, and Search Engine Land documents the competitor-analysis version of it in detail (entity-based competitor analysis guide).

Do not treat the diff as a checklist to stuff. A missing entity is a question, not a task. It asks whether the competitor covers something your page genuinely should, or whether they are simply broader in scope. The audit surfaces the candidate. Judgment, yours, decides what to write.

## Where an agentic coding tool earns its place

The four steps above are glue: fetch, call an API, join, diff. An agentic coding tool like Claude Code, Codex, or Antigravity is well suited to writing and maintaining that glue, because the shape of it changes with every site. Your schema lives in a different place, your renderer behaves differently, your competitor set rotates. The agent regenerates the extraction and join logic against those specifics instead of you maintaining a brittle scraper by hand.

The more useful role is interpretation. Once the diff table exists, an agent can read each gap in context, the way it reads a quality framework rather than a scoring endpoint, an approach I walked through when building an E-E-A-T checker with AI. It can note that an undeclared high-salience entity has an obvious sameAs target, or that a declared entity never appears in the prose and so was never going to be recognized. That reasoning over the gap is the part a query alone cannot give you.

## The caveats that keep this honest

Three failure modes will bite you, so name them before you present results.

- The Natural Language API is not Google Search’s ranking Knowledge Graph. It is a documented proxy, and its coverage of MIDs is partial, so absence of a MID is not proof Google fails to recognize an entity.
- Salience is relative to the document, not to your site’s authority for that entity. A high score means “central to this page,” never “you own this topic.”
- Adding a sameAs does not force an entity into the Knowledge Graph. Structured data is a strong hint, not a write API, and the graph updates on its own schedule.

Hold those three in view and the audit stays defensible. Drop them and you will over-claim, presenting a proxy measurement as if it were Google’s own verdict. If you want a starting point, wire up steps one and two against a single page this week, look at the two lists side by side, and let the size of the gap tell you whether the full pipeline is worth building.

    Tags [#entity-gap-audit](/tags/entity-gap-audit)[#entity-seo](/tags/entity-seo)[#schema-org](/tags/schema-org)[#google-nlp-api](/tags/google-nlp-api)[#agentic-coding-tools](/tags/agentic-coding-tools)   Share        Written by [Carlos Arias](/authors/carlos-arias)

Builder of AstroAgent, an AI-run website platform.

         On this page

- What an entity gap audit actually measures
- Step one: extract the entities you declare
- Step two: extract the entities Google’s NLP recognizes
- Step three: join the two sets into a queryable graph
- Step four: add the competitor diff
- Where an agentic coding tool earns its place
- The caveats that keep this honest

## Continue reading

      [Guides](/categories/guides) · September 30, 2026  [### Meta Ads Creative Diversity: Why Variety Beats Volume](/blog/guides/meta-ads-creative-diversity-over-volume)

Meta ads creative diversity now outranks volume. Why varied accounts still rate Low, plus a ranked checklist for hooks, creators, formats and offers.

  Carlos Arias · 5 min
      [Guides](/categories/guides) · September 26, 2026  [### SEO Content Conversion After the Click: A Build Guide](/blog/guides/seo-content-that-converts-after-the-click)

SEO content conversion after the click is decided by the page you engineer, not the traffic handed to sales. How to build the path.

  Carlos Arias · 6 min
      [Guides](/categories/guides) · September 25, 2026  [### AI Coding Assistant Domain Knowledge: Grounding in Schemas](/blog/guides/grounding-ai-coding-assistants-domain-knowledge)

AI coding assistant domain knowledge fails when it is memorized. Google's Ads API Developer Assistant v4.0.0 shows why schema grounding beats guessing.

  Carlos Arias · 5 min

## Stay in the loop.

One email when it’s worth it — new posts and updates, no spam.

Thanks — check your inbox to confirm.

Free. Unsubscribe in one click.
