VisibilityReport

Guide · September 2026

Knowledge graphs and Wikidata: how machines learn who you are, and why coverage comes first

A knowledge graph is a machine’s answer to the question “who is this, exactly?” It stores brands, people and products as connected entities with verified facts, and AI assistants lean on it when they decide whether the name in a forum thread, a news article or a buying question is you, a competitor, or a namesake on another continent.

Unlike most technical visibility work, you cannot simply implement your way into it. Entry runs through notability: independent sources writing about you. This article explains how knowledge graphs and Wikidata work, what notability actually requires, how to earn the coverage that unlocks it, and why one entity serves every language and market you sell in.

What is a knowledge graph?

A knowledge graph is a database of entities: things, not strings. Where a text index stores the word “jaguar” and hopes context sorts out cat versus car brand, a knowledge graph stores distinct entities with typed facts and relationships: Jaguar (car manufacturer), founded 1922, headquartered in Coventry, subsidiary of Tata Motors.

Google introduced its Knowledge Graph in 2012 under exactly that slogan, “things, not strings,” and it has quietly become core infrastructure. It powers the knowledge panels beside search results, but its more important job is invisible: disambiguation and trust. When a system reads “Example Store raised prices,” the graph is how it knows which of the world’s Example Stores that sentence is about, and which established facts the claim can be checked against.

That job transfers directly to AI search. Assistants like ChatGPT, Claude, Gemini and Perplexity assemble answers from many sources, and entity data is part of how brand mentions get resolved, connected and trusted. A brand that exists as a clean entity is a brand machines can confidently talk about.

What is Wikidata?

Wikidata is the world’s largest open knowledge graph: a free, collaboratively edited database run by the Wikimedia movement, launched in 2012, holding well over 100 million items. Every item has a permanent identifier, a Q-number like Q95 for Google, plus labels and descriptions in any language, statements (“instance of: business”, “inception: 2016”, “official website: …”), and references backing them.

Three properties make it the entity hub rather than just another wiki. It is radically multilingual: one item carries names and descriptions in every language, so the same identity serves a customer in Oslo and an AI answering in Spanish. It is freely reusable: the data is published under an open license, which is why it flows into search engines, Wikipedia’s infoboxes, voice assistants and countless datasets, including the corpora AI models train on. And it is referenced by design: statements cite sources, which is what makes machines treat it as checkable rather than promotional.

Its role in the AI stack is deepening. In October 2025, Wikimedia Deutschland launched the Wikidata Embedding Project: a vector-based semantic search over the whole graph, built for AI applications and supporting the Model Context Protocol (MCP), the emerging standard for connecting AI systems to data sources. Wikidata is no longer just something models trained on; it is being wired in as live, verifiable grounding.

Why do knowledge graphs matter for AI visibility?

Because before an AI system can recommend you, it has to be sure who you are. Entity data is where that certainty comes from, and it does three concrete jobs: it disambiguates your name from every namesake, it binds your scattered mentions (press, reviews, forums) to one identity, and it gives generated claims about you a factual backbone.

Our August 2026 audit of 99 Norwegian online stores (in Norwegian) measured this layer, and the gap is wide: 47 of 99 stores had no Wikidata entity at all. The distribution is telling. Among pure-play niche stores, 85% lacked an entity; among chain-owned stores, 12.5%. Yet three of the smallest, least-authoritative stores in the study had entities as complete as the market’s biggest names, proof that this is not gated by size or budget.

Honesty requires the counterweight, and our data provides it: among low-authority stores, a complete entity alone produced no measurable lift in AI citations. One pair in the study makes the point from the other side: two sister chains on technically identical platforms, where the one with the Wikipedia-connected identity ranks on its own local searches and the one without does not rank at all. The entity is not a lever you pull for citations. It is hygiene: the identity layer that makes every mention you earn attributable to you. Which raises the real question: how do you earn the mentions?

What is notability, and why is it the gatekeeper?

Notability is the requirement that independent, serious sources have described you. It is the admission ticket to the whole graph layer, and the bar differs sharply between the two systems people conflate:

Wikipedia’s bar is high. An article requires significant coverage in reliable sources that are independent of the subject: real journalism or literature about you, not press releases, not your own site, not directory listings. Most small and mid-sized businesses genuinely do not clear it, and articles that try anyway get deleted.

Wikidata’s bar is lower, but real. An item qualifies if it meets any of three criteria: it has a valid page on a Wikimedia project (a Wikipedia article, for instance); or it refers to a clearly identifiable entity that can be described using serious and publicly available references; or it fills a structural need in the graph. For a registered company, criterion two is the usual door: a company-register entry, trade-press coverage, product reviews in credible publications. No Wikipedia article required.

What notability is not: something you can purchase or mark up. Advertorials, bought placements and self-published content do not count as independent sources, and promotional items without genuine references get flagged and deleted by the community. The gate exists precisely so the graph stays trustworthy enough for machines to rely on. That is the deal: the graph is valuable to you because you cannot spam your way in.

How do you earn the coverage that makes you notable?

By giving independent sources a reason to write about you, which is the same work that drives AI recommendations directly, since assistants consult that coverage when answering buying questions. The graph entry just binds it to your name. Ordered by return on effort:

  1. Be the obvious answer to something. Coverage follows distinctiveness. “Another online store with a wide selection” gives a journalist nothing; “the specialist for tactical outdoor gear” or “the only carbon-neutral furniture shop in the region” is a story hook and a quotable label. Sharpen the one-line answer to “what are you alone in?” before pitching anyone. Our SEO, AEO, GEO and AIO guide covers this kind of answer-shaped positioning in more depth.
  2. Publish original data. Surveys, price analyses, industry statistics from your own operations: original numbers are the most reliably cited asset in existence, because every journalist and every AI answer needs sources for claims. One genuine dataset outperforms a year of press releases.
  3. Court the trade and local press first. Industry publications and regional media have lower thresholds than national outlets, and they count fully as independent sources, for Wikidata, and for the AI models reading them. A profile in a retail-trade journal is worth more than ten self-published posts.
  4. Get your products independently tested. Review sites, comparison portals, category “best of” lists: these are precisely the sources AI assistants consult on buying questions, and each one is citable coverage.
  5. Lend expertise to journalists. Journalist-request platforms and direct outreach make your founders quotable experts. Every quoted appearance is an independent mention with your name attached to your specialty.
  6. Collect the institutional signals. Industry awards, association memberships, chamber listings, conference talks, podcast appearances. Individually small; together they form the reference base that “serious and publicly available” points to.

Two honest notes. This track takes months to years: it is the slowest asset in online visibility, which is why starting early beats starting big. And it cannot be delegated to markup: no structured data substitutes for a source that chose, independently, to describe you.

How do you create and maintain a Wikidata entity, step by step?

Once genuine sources exist, the entity itself is hours of work, not months:

  1. Check whether you already exist. Search Wikidata for your brand and prior names. An unclaimed or wrong item is more urgent than a missing one: fix before you create.
  2. Gather references first. Company-register entry, official website, and the independent coverage you have earned. Every statement you add should be backed by one of these.
  3. Create the item neutrally. Label (brand name) and a short, factual description, such as “Norwegian online retailer of outdoor equipment” rather than marketing language, in your own language and English at minimum. Descriptions also disambiguate you from every same-named entity on earth.
  4. Add the core statements. Instance of (business / online store), inception date, country, headquarters location, industry, official website, founder or parent organization where relevant, and a logo via Wikimedia Commons, noting that Commons requires the logo be freely licensed or below the threshold of originality.
  5. Add identifiers, the strongest disambiguation you can buy for free. Your national company-register ID (nearly every country has one: Brønnøysund, Companies House, Handelsregister and their siblings), plus official social profiles. Identifiers are what let machines join your entity to authoritative records.
  6. Reference every substantive statement. Items with sources survive community review; bare assertions invite deletion debates.
  7. Close the loop from your own site. Point the sameAs property in your Organization JSON-LD at your new Q-identifier, so your domain and your entity confirm each other. Our JSON-LD article covers that block in detail.
  8. Maintain it. Put the item on a watchlist and update facts when they change. Editing items you are affiliated with is acceptable on Wikidata when edits are factual and referenced: the etiquette is neutrality, not absence. Writing your own Wikipedia article is a different matter: conflict-of-interest rules are strict, undisclosed paid editing violates the terms of use, and the practical advice is simple. Let a Wikipedia article happen to you when the coverage justifies it.

What does this look like in a global context?

One entity, every market: that is the quiet superpower of this layer. A Wikidata item is not a national record: the same Q-identifier carries your Norwegian description, your German description and your Japanese one, and the AI assistant answering a shopper in any of those languages resolves you to the same identity. For anyone selling across borders, translating your entity’s labels and descriptions is the cheapest international visibility work that exists, minutes per language.

The global frame also explains why the layer matters more, not less, as you grow. Name collisions are a global problem: the bigger the world you operate in, the more same-named companies, products and people you compete with for machine understanding. Disambiguation is what entity data is for. And the infrastructure is consolidating around open, multilingual backbones: Google’s graph draws on Wikidata among its sources, Wikipedia’s infoboxes draw from it, LLM training corpora include it, and the Embedding Project now serves it directly into AI applications across a hundred-plus languages. Meanwhile the reference base underneath (company registers, trade press, review sites) exists in some national form virtually everywhere, so the playbook above translates: register, earn coverage in the sources your market trusts, encode the identity, repeat per market only where the coverage itself must be local.

Frequently asked questions

Can I create a Wikidata item for my own company?

Yes, if the notability criteria are met. Affiliated editing is acceptable on Wikidata when statements are neutral, factual and referenced. The item must describe a clearly identifiable entity backed by serious public sources, such as a company register and independent coverage. Purely promotional items without references are deleted by the community.

Do I need a Wikipedia article to be in the Knowledge Graph?

No. Wikipedia’s notability bar (significant independent coverage) is far higher than Wikidata’s, and most businesses don’t clear it. A well-referenced Wikidata item, a company-register entry and consistent Organization markup already give machines a verifiable identity. Treat a Wikipedia article as a possible outcome of years of coverage, not a prerequisite.

Is a knowledge panel the same as being in the Knowledge Graph?

No. The panel is the visible tip. Google’s Knowledge Graph contains vastly more entities than ever get panels, and panel display varies by market and query. Your goal is the entity layer itself: being unambiguous and verifiable to machines. Whether a panel appears is Google’s choice, and no one can guarantee it.

Will a Wikidata entity make AI assistants cite my store more?

Not by itself. In our study of 99 online stores, complete entities alone showed no measurable citation advantage. The entity makes you unambiguous; the independent coverage behind it is what AI systems actually consult when recommending. Do both, in that order: earn the mentions, then encode the identity.

How long does this take?

The entity work is fast: with sources in hand, a complete, referenced Wikidata item is a few hours. The coverage that justifies it is the slow part: realistic timelines for meaningful independent mentions run months to years, which is why reputation work should start alongside, not after, your technical fixes.

Final thoughts

Of the 99 stores in our study, 47 are missing from the world’s open identity layer entirely, and among the small, specialized stores that AI search otherwise favors, it is 85%. That is the pattern worth sitting with: the players with the most to gain from being unambiguous are the least likely to have done the work, while the work itself (a referenced entity on top of earned coverage) is gated by neither size nor budget.

But keep the order straight, because it is the whole lesson. Coverage first: be distinctly something, publish facts worth citing, and let independent sources say so. Identity second: encode what the sources establish into Wikidata, your markup, your registers. Machines don’t recommend you because you have an entity. They recommend you because trustworthy sources describe you, and the entity is how they know, in every language at once, that those sources mean you.