← Back to blog
Published · Updated

Signal vs hype: what 414 AI terms reveal about operational maturity

by Skills Intelligence

An enterprise AI meeting can now pass without anybody using a stable noun.

A supplier promises an agentic platform. A strategy deck calls the operating model AI-native. Learning proposes a course in reasoning. Procurement asks for a copilot. Everyone recognises the words, so the conversation feels precise. Then ask what the system may do, how its capability will be tested, or which failure remains a human responsibility. The apparent agreement disappears.

This is not pedantry. An unstable term can become a product requirement, a budget line, a skill in a taxonomy, and eventually a criterion applied to people. By then, the word is doing the work that evidence should have done.

The 414-term AI glossary on this site offers an unusually useful view of that problem. It also sharpens a distinction that the dataset itself cannot measure: popularity and operational maturity are different things.

Repetition creates attention. A term becomes operational only when somebody attaches consequences to a definition: an interface to implement, a test to pass, an obligation to meet, or a procedure to follow.

That distinction matters to anyone buying technology or building skills intelligence. Hype can be a valuable early signal. It is a very poor data model.

What this dataset can — and cannot — tell us

The glossary was assembled across three editorial waves from 22 AI-assisted source runs, plus 17 externally sourced entries, then consolidated and connected to papers, standards, laws, institutional publications, and first-party documentation where available. These are runs, not 22 statistically independent models.

The figures below use the resolved catalogue served by /glossary, version 2026-06-26+editorial-glossary-2026-01, verified on 28 August 2026 — not the unreviewed base file. It contains 414 terms and 448 term–reference records pointing to 415 unique URLs. Of the 414 entries, 320 have at least one reference and 94 have none.

This source process is not identical to the three-run skills-atlas experiment documented in the methodology article. The transferable rule is narrower: model agreement is provenance about a research procedure, not proof of the claim on which the models agree.

Each entry also has an editorial maturity score:

LevelWorking interpretationTermsShare
1Neologism: a new or narrowly sourced label10425%
2Buzzword: visible, but still contested or imprecise12731%
3Technical term: used with a reasonably specific meaning11127%
4Standardised: anchored in established documentation or practice4912%
5Canonical: formally codified or institutionally settled236%

That puts 231 terms, or 56%, at levels 1 and 2. Only 72, or 17%, sit at levels 4 and 5. A separate flag marks 56 terms as speculative.

Those numbers describe this collection, not the whole AI discourse. The sample was built to map an expanding frontier, so it deliberately over-represents new language. It has no search-volume, citation-frequency, or longitudinal usage series, so it cannot measure a term’s popularity or prove that attention is rising. The score is an editorial judgement, not a measured property of a term. Reference count is not evidence quality. Nor does maturity mean importance, technical merit, adoption, or legal force.

So “56% froth” is a useful provocation, but it does not mean that 56% is nonsense. Low-maturity terms can provide leads for horizon scanning; this collection does not quantify the attention behind them. Some of today’s unsettled phrases will become tomorrow’s useful categories. Others will disappear. The management error is not to notice them; it is to treat them as settled before they have earned that status.

For two different kinds of emerging language, compare the LLM OS metaphor with the mHC research method. The glossary entries explain what each name refers to and where its evidence stops; neither should be read as a certificate of production readiness.

Formalisation beats fame

The working thesis of this essay is that terms become usable when they acquire one or more forms of institutional support. “Institutional” is broader than “regulatory”. Four mechanisms matter, and they should not be confused.

1. Law makes definitions consequential

Article 3 of the consolidated EU AI Act defines terms including AI system, provider, deployer, and general-purpose AI model. The linked consolidation is current to 27 July 2026; Article 113 sets 2 August 2026 as the general date of application while preserving a staggered timetable for specified provisions. These definitions are not philosophically final. They are operational because their scope changes who has which obligation.

This is a particular kind of maturity: legal scope. A legally defined term can remain technically controversial, and a popular technical term can remain absent from law. In an enterprise glossary, “defined in a regulation” should therefore be a source status, not a synonym for “universally true”. Jurisdiction, entry into force, date of application, and relevant transition belong beside the definition.

2. Shared vocabularies make coordination cheaper

ISO/IEC 22989:2022 establishes AI terminology to support communication among different stakeholders and the development of other standards. It is an international standard. The NIST trustworthy AI glossary is an institutional publication, not an ISO standard; it has a related but revealing design choice. It preserves multiple meanings used by different disciplines instead of pretending that one sentence has ended every debate.

The OECD framework for classifying AI systems is a policy-classification framework. It organises systems across People & Planet, Economic Context, Data & Input, AI Model, and Task & Output so that technical characteristics can be connected to policy implications, inventories, and risk assessment.

These artefacts do not have the same authority, and none makes language permanent. They make disagreement more inspectable by giving a term a source, scope, version, and relationship to a decision. That is enough to coordinate work without demanding metaphysical consensus.

3. Protocols turn nouns into interfaces

Consider MCP, one of the more connected terms in this collection. The Model Context Protocol specification dated 28 July 2026 does not merely describe a fashion. It identifies hosts, clients, and servers; defines JSON-RPC messages, stateless requests, per-request capability negotiation, resources, tools, and prompts; and uses normative requirements. Implementers can test whether two components exchange the expected messages. This version also changed the core session model, which is a useful reminder that a protocol’s version is not a footnote.

That makes the term more operational than “agentic AI”. It does not make every MCP implementation interoperable, secure, or valuable. A protocol name is the start of a technical contract, not evidence that a supplier has fulfilled it. Version, supported capabilities, authorisation, observability, and conformance still have to be examined.

4. Operating policies make internal thresholds explicit

Anthropic’s Responsible Scaling Policy and its AI Safety Levels provide an example of corporate operating language. The current RSP materials connect capability thresholds to safeguards and internal decisions. That can make the terms meaningful inside a governance system even though the policy is not a statute.

Similarly, the European Commission describes the General-Purpose AI Code of Practice as an adequate voluntary tool for providers to demonstrate compliance with binding AI Act obligations. Signing the code is voluntary; the underlying obligations are not. “Voluntary code”, “company policy”, “technical standard”, and “law” are four different statuses. Collapsing them into a single maturity label throws away exactly the information an executive needs.

A correction that proves the point

An earlier version of this article listed California SB-1047 among measures “written into law”. That was wrong. Governor Gavin Newsom returned the bill without his signature in September 2024.

The bill still mattered. It influenced the policy debate and the vocabulary of frontier model governance. But salience is not enactment, a bill is not a statute, and a citation to legislative text is not proof that the text entered into force.

California later enacted a different measure: SB-53, the Transparency in Frontier Artificial Intelligence Act, which Governor Newsom signed on 29 September 2025. The sequence is another reason to identify the exact instrument, status, and date rather than rely on a general regulatory storyline.

This is more than a correction to one sentence. It exposes a weakness of every maturity scale, including ours: one number can conceal the type and status of the underlying authority. A reliable terminology record needs at least a source type, jurisdiction or domain, version, status, verification date, and named editorial owner. The maturity score may help people filter; it must never replace provenance.

The “hubs” are routes through this glossary, not laws of the field

The resolved glossary records explicit editorial relationships between selected terms. Thirteen terms tie at five relationships: RLHF, DPO, RLVR, model collapse, synthetic data, knowledge distillation, computer use, prompt caching, Constitutional AI, LLM evaluations, LLM-as-a-judge, context engineering, and Agent2Agent Protocol. Another 35 have four, including MCP, reasoning models, test-time compute, RAG, AI slop, and LLM Wiki.

That is useful for navigation. It is not a bibliometric network analysis, and it does not prove that these are the universally “load-bearing” ideas of AI. The presence of both technical mechanisms and cultural labels in the leading group is a warning: relationship count partly reflects the editorial map we chose to draw.

A better learning sequence begins with the decision a person needs to make. Someone evaluating inference performance may need scaling laws, post-training, test-time compute, and evaluation design. Someone governing tool-using systems may need permissions, identity, protocol architecture, logging, and human control before learning the latest agent label. A graph can expose dependencies, as the article on learning prerequisite graphs argues, but the graph must be built for a purpose.

Use a term-state model, not one universal maturity score

For enterprise use, replace the single ladder with a small state record. A term may occupy several states at once.

StateThe question it answersAppropriate use
Attention signalAre credible people beginning to use it?Horizon scanning and research backlog
Working termDo we have a scoped internal definition?Discussion, discovery, and provisional tagging
Technical contractIs there a versioned interface, test, or method?Architecture, engineering, and assurance
Controlled termIs it owned, governed, and mapped to evidence?Taxonomies, assessments, and workflows
Regulated termDoes a binding legal instrument define it for a jurisdiction and application date?Compliance analysis with legal review

This avoids a false choice between ignoring hype and institutionalising it. “Agentic AI” can be a strong attention signal and a weak controlled term at the same time. “AI system” can be a regulated term under one legal instrument while carrying a different scope in a technical standard. The record preserves both facts.

A six-question test before a buzzword enters a skills system

Before adding an AI term to a capability model, procurement requirement, or learning catalogue, ask six questions.

  1. What observable difference does the term describe? Name the task, behaviour, or decision that changes. “Understands agentic AI” is not observable.
  2. Which definition are we adopting? Record its source, status, version or date, and important exclusions. Do not silently blend a vendor definition with a legal one.
  3. What is the unit? Is this a technology, capability, method, role, product category, risk, or cultural label? A taxonomy fails when these become siblings without types.
  4. What evidence would support a claim? A course completion, product demo, code sample, scenario assessment, and production outcome support very different claims.
  5. Which decision may use it? Discovery tolerates ambiguity that selection, authorisation, or redundancy decisions should not.
  6. Who reviews and retires it? Emerging language needs an owner, review date, correction path, and downstream impact check.

If a term cannot pass the first three questions, keep it on the watchlist. If it cannot pass all six, do not let it become a high-impact decision criterion.

What this changes in procurement, learning, and workforce planning

In procurement, translate category claims into testable statements. If a supplier offers “agentic workflows”, ask which tasks the system initiates, which tools it may call, how permissions are bounded, what is logged, where approval is required, and which failures the acceptance test covers. A fashionable noun should never be accepted as a feature.

In learning, teach stable primitives before transient labels. Employees need to understand evaluation, evidence, context, permissions, failure modes, and task redesign. They can then interpret the next wave of terminology without requiring a new course for every slogan. Where a new term matters, assess an observable capability rather than recall of the definition.

In workforce planning, do not infer a workforce gap from search volume. Rising use of a phrase may justify investigation; it does not tell you how many people need which capability, at what level, by when. Connect the term to changing tasks, operating demand, evidence standards, and a decision owner. The distinction between inference and evidence is developed further in AI skills intelligence.

And in governance, preserve the unpopular metadata. Source type, status, version, scope, provenance, and correction history are less exciting than a trend chart. They are also what stops yesterday’s marketing language from becoming tomorrow’s organisational fact. The article on skills intelligence governance sets out the corresponding decision rights.

The real signal inside the hype

The distribution cannot establish a causal law about how vocabulary settles. Read beside the examples above, it supports a working hypothesis: language becomes operational when institutions make it costly to remain vague. Laws assign obligations. Standards coordinate meanings. Protocols expose incompatible implementations. Operating policies connect thresholds to actions. Tests allow claims to fail.

Popularity does none of those things on its own.

So monitor low-maturity language as a source of hypotheses. Do not pour it directly into a taxonomy, a contract, or a career decision. First turn the word into a scoped claim, the claim into evidence, and the evidence into a decision with an owner.

If your organisation needs to convert fast-moving AI language into governed capabilities, evaluation criteria, and an implementable skills model, that boundary between signal and operating truth is a good place to start the work.

Continue exploring

— Skills Intelligence