← Back to blog
Published

What is skills intelligence? A practical model for AI work

by Skills Intelligence

Imagine a dashboard reporting that 312 employees have "cloud architecture". Before acting on it, ask four questions. Where did the claim come from? How strongly does the evidence support it? When was it last true? What decision is it supposed to inform?

Perhaps an employee selected the label in a profile three years ago. Perhaps a model inferred it from a job title. Perhaps it comes from a recent assessment and two delivered projects. All three records may display the same green tick, while supporting radically different decisions. The self-report may be enough to suggest a course. It is not enough to appoint the technical authority for a high-risk migration.

This is the central problem skills intelligence should solve. It is not a database of skill names, and it is not an AI product that somehow "knows" the workforce. Skills intelligence is the discipline of turning evidence about capability and work into decisions, while preserving the limits of that evidence.

That definition extends a well-established public-policy tradition. Cedefop defines skills intelligence as the outcome of an expert-driven process of identifying, analysing, synthesising, and presenting quantitative or qualitative information about skills and labour markets. At enterprise level, the same logic applies to hiring, learning, mobility, staffing, role design, and workforce planning. Data becomes intelligence only when people can use it for a defined purpose and examine how it was produced.

The category error: buying a noun instead of building a capability

The market uses several related terms as if they were interchangeable. They are not.

TermWhat it answersWhat it cannot answer by itself
Skills inventoryWhich skills have been recorded against people or roles?Whether those records are comparable, current, or evidenced.
Skills taxonomyHow are skills named and grouped?Whether someone possesses a skill or a role genuinely requires it.
Skills ontology or graphHow do skills, tasks, roles, evidence, and resources relate?Whether an edge is true enough for a particular use.
Skills analyticsWhat patterns appear in the available data?Whether missing or biased data makes the pattern misleading.
Skills platformHow can data and workflows be operated at scale?Whether the organisation has sound decision rules or governance.
Skills intelligenceWhat does the evidence justify doing, for whom, and with what uncertainty?Accountable human judgment; the organisation still owns the decision.

A taxonomy is valuable infrastructure. The OECD's 2026 work on a common skills language describes shared definitions as foundational to skills-first systems, while also identifying granularity, duplication, interoperability, and continuous governance as real design problems. ESCO, for example, supplies stable identifiers, descriptions, relationships, and multilingual reference data that applications can reuse. But a reference language does not know whether a particular engineer can design a production system, or whether that ability should outweigh availability, aspiration, or domain experience when staffing a project. Even an apparently simple label such as prompt engineering needs scope: casual interaction with a model is not the same capability as designing and evaluating prompts inside a production workflow.

That last mile is not a feature to purchase. It is an operating capability to design.

The four-coordinate rule

Every claim about a skill should travel with four coordinates: provenance, confidence, recency, and decision context. Call this the evidence envelope. A claim without its envelope may still be useful for search or exploration; it is not decision-grade intelligence.

CoordinateQuestion it must answerWhat commonly goes wrong
ProvenanceWhat exactly was observed, where, by whom, and under which permissions?A model output looks identical to an assessment or verified work product.
ConfidenceHow strongly does the available evidence support this exact claim?Confidence is confused with proficiency, or an unexplained score creates false precision.
RecencyWhen was the evidence produced, and how quickly might it decay?Old records remain "current" because nothing triggers review.
Decision contextWhich action may use the claim, at what stakes, and with whose review?Evidence collected for development quietly becomes a gate for employment.

1. Provenance: preserve the chain of reasoning

"Python" on a CV is an observed text mention. "Likely knows Python" is an inference. Passing a job-relevant assessment is evidence under specified conditions. Delivering a maintained Python service is evidence of applied work, although even that says little about every Python task.

These signals can reinforce one another, but they should never be silently collapsed. A useful record retains the source type, date, method, transformation, owner, and permission for use. If a model mapped "data scripting" to Python, the original text and mapping rule matter. If a manager validated a level, the rubric and scope matter. Provenance turns disagreement from a trust crisis into an inspectable question.

2. Confidence: do not launder inference into fact

Confidence answers how well the evidence supports a claim. Proficiency answers how well someone can perform. They are separate dimensions. A person can have high proficiency supported by weak evidence, or a confidently observed beginner level.

AI makes this distinction easy to lose. A model can classify text, reconcile synonyms, or propose relationships at useful scale. Its output remains probabilistic. A percentage deserves to be shown only if it has been evaluated and calibrated for the relevant population and use; otherwise, plain-language bands with explicit criteria may be more honest. A polished score does not become reliable merely because the interface omits the uncertainty. The practical safeguards are covered in depth in AI skills intelligence: why inference is not evidence.

3. Recency: every skills record has a clock

Freshness is not one universal expiry date. A safety certification has a formal validity period. A cloud API can change in months. Negotiation experience may remain useful for years while its context changes. Emerging terminology can be renamed before an annual profile review.

The system therefore needs a recency policy by evidence and skill type: timestamp, review trigger, decay rule where appropriate, and a way to distinguish "old" from "disproved". Cedefop explicitly describes skills intelligence as continuous and iterative. A static inventory may document the past accurately and still recommend the wrong future.

4. Decision context: adequacy depends on consequence

There is no context-free threshold for "good enough" skills data. The evidence needed to recommend an optional learning module is different from the evidence needed to reject an applicant, set pay, allocate safety-critical work, or select people for redundancy.

This principle aligns with the NIST AI Risk Management Framework, which treats validity, transparency, fairness, and other trustworthiness characteristics in the system's context of use. It is also more than prudent design. The EU AI Act classifies specified AI systems intended for recruitment, selection, promotion or termination, task allocation based on personal characteristics, and worker monitoring or evaluation as high-risk, subject to the Regulation's scope and conditions. A learning suggestion and a promotion decision may use similar data; they are not the same system in risk terms.

One claim, three evidence thresholds

The four coordinates prevent an organisation from applying one universal "truth" about a person to every workflow.

UseA proportionate starting pointWhat the output should not claim
Discovery: suggest learning, mentors, or opportunitiesRecent self-report or clearly labelled inference, visible to and correctable by the personVerified proficiency or eligibility
Planning: estimate capability supply for a team or job familyMultiple signals, known coverage, aggregation, domain review, and explicit uncertaintyA complete census of individual capability
Consequential selection: hiring, promotion, pay, or sensitive work allocationCurrent, job-relevant evidence; documented criteria; validation for the population; human accountability and a challenge pathThat an opaque match score is the decision

This is why "single source of truth" is often the wrong ambition. Skills intelligence needs a governed source of claims and evidence, not a machine that turns all ambiguity into one canonical answer. Contradiction is information: an assessment, self-report, and work history may disagree because they measure different things.

The decision loop: frame, observe, interpret, act, learn

The evidence envelope makes individual claims safer. A five-stage loop makes the capability useful.

  1. Frame one decision. Name the action, affected population, owner, baseline, time horizon, and cost of error. "Gain visibility" is not a decision. "Decide whether to build, hire, or contract for cloud security capacity in the next two quarters" is.
  2. Observe within a declared boundary. Choose evidence that is relevant and permitted. Record which teams, languages, roles, systems, and time periods are represented—and which are absent.
  3. Interpret with a shared language. Reconcile aliases, map work to skills, represent useful relationships, and separate observations from inferences and editorial judgment.
  4. Act through an accountable workflow. Put the result into workforce planning, staffing, learning, or hiring with explicit decision rights, review, and correction.
  5. Learn from the outcome. Did the intervention improve time to staff, internal fill, verified capability, delivery quality, or another business measure? Corrections and outcomes should update the model rather than disappear into a dashboard.

The order matters. Starting with an enterprise catalogue reverses it: the team builds an elaborate answer before agreeing on the question. Catalogue size then becomes a proxy for progress. In reality, every additional concept creates definitions, mappings, evidence rules, ownership, and maintenance work. A catalogue can grow faster than the organisation's ability to know what any of it means.

A practical example: the cloud transformation question

Consider a fictional but typical organisation preparing to move several regulated services to the cloud. Leadership asks, "Do we have the skills?" A weak implementation searches profiles for "cloud", counts matches, and colours a dashboard.

A decision-led implementation asks what choices are pending. Which services will be redesigned? Which work requires architecture, security, platform engineering, vendor management, or regulated operations? What level is needed, by when, and where would an error be costly?

It can then combine role and project history, relevant credentials, work samples or assessments, employee aspirations, and expert validation. Each claim keeps its evidence envelope. The result is not a ranking of "cloud-ready people". It is a set of bounded conclusions: where capability is sufficiently evidenced, where a short validation is needed, where adjacent skills make internal development plausible, and where external capacity is the defensible choice.

That output can change a build-buy-borrow decision. The keyword count cannot.

What serious implementation looks like

The technical architecture still matters. Organisations need stable concepts, identity and access controls, source integration, relationships among skills and work, evaluation, correction, and versioning. Platforms can make those activities operable at enterprise scale. Our companion guide explains how to evaluate a skills intelligence platform.

But the first implementation unit should be a decision, not the enterprise. A credible starting sequence is:

  1. select one business decision and one bounded population;
  2. write the decision contract: owner, permitted uses, evidence threshold, harms, and success measure;
  3. define a small governed vocabulary and the four-coordinate evidence model;
  4. test concepts and inferences with domain experts and affected people;
  5. run the workflow against a baseline, including sparse and contradictory cases;
  6. inspect errors and distribution of outcomes, then decide whether to scale;
  7. version the model and publish a correction path.

This is slower than a compelling demo and faster than an enterprise taxonomy programme. More importantly, it can fail informatively. A pilot that reveals insufficient evidence, poor job architecture, or an unusable workflow has created value before those weaknesses are scaled.

Four forms of skills theatre

  • Taxonomy theatre: success is measured by the number of concepts approved, not decisions improved.
  • Inference laundering: a possible skill enters the interface and emerges as a fact because its origin is no longer visible.
  • Real-time theatre: continually refreshed signals are presented as a live map of human capability. There is no real-time view of a person's complete capability—only evidence arriving at different speeds.
  • Dashboard theatre: profile completion, logins, and visual coverage rise while hiring, mobility, learning, staffing, or planning remains unchanged.

These are not cosmetic problems. Each one converts uncertainty into organisational confidence without adding evidence. The remedy is not necessarily another tool. It is a better decision contract, evidence model, validation process, and operating owner.

How this project applies the model

Skills Intelligence is an independent research project about technical work and AI. It is not an employee-ranking system or a commercial talent platform. The current public dataset includes:

  • 426 technical skills organized into 15 domains.
  • 292 prerequisite relations for learning order.
  • 78 professional roles assessed by skill-level AI pressure.
  • 400 AI terms tracked by origin and maturity.

The AI Skills Atlas supplies defined concepts and prerequisite relationships. The role dictionary organises work from observed skill profiles. The AI Exposure research explores directional pressure at skill and role level, while the AI glossary tracks a fast-changing vocabulary.

The project also demonstrates the limits of the approach. Research-system agreement is shown as provenance, not treated as a quality score. Prerequisites and exposure verdicts contain editorial judgment. Coverage is selective. The data is versioned research, not ground truth, and it should not be used for automated employment decisions. The methodology and correction policy make those boundaries visible.

The shortest useful definition

Skills intelligence is a decision discipline built on a shared language and inspectable evidence. Its unit of value is not the skill record, graph edge, model score, or dashboard. It is a better, more defensible action—and the capacity to learn when the evidence was wrong.

If a system cannot show provenance, confidence, recency, and decision context, it may contain skills data. It does not yet contain skills intelligence.

If you need to turn a workforce question into a decision contract, evidence model, or bounded pilot, get in touch to discuss the implementation.

Continue exploring