← Back to blog
Published

AI skills intelligence: why inference is not evidence

by Skills Intelligence

The most dangerous sentence in a skills-platform demonstration is not “our AI inferred 80% of your workforce’s skills.” It is the sentence that often follows: “your skills data is now complete.”

AI can turn job titles, CVs, project histories, learning records, and work artefacts into a plausible skills profile. It can replace a blank page with candidates worth reviewing. But it has not observed proficiency. It has produced hypotheses from traces.

This is the line between useful intelligence and automated organisational fiction:

A skill inference is a claim to investigate, not evidence that a person can perform.

The central design problem in AI skills intelligence is therefore not how to infer more. It is how to preserve the boundary between a useful hypothesis and a claim that has earned the right to influence a decision.

If you are new to the wider discipline, start with what skills intelligence is. Here we will focus on the harder implementation question: how to use inference without laundering uncertainty into employment decisions.

What a model actually knows

Imagine an engineer whose project update says, “helped migrate reporting workloads from PostgreSQL.” A model could reasonably extract PostgreSQL and database migration. It might also infer SQL, data modelling, cloud migration, performance tuning, or technical leadership. Each step moves further from the observed sentence.

None establishes whether the engineer designed the migration, reviewed one query, managed the team, or attended the meeting—let alone proficiency, recency, or independence.

The same problem appears across common sources:

SignalWhat was observedA defensible claimWhat it does not prove
Skill on a CVThe person chose a labelSelf-reported exposure or experienceCurrent proficiency or independent performance
Job titleThe person occupied a rolePossible role-related skillsThat this job matched the standard role profile
Course completionA learning event was completedExposure to defined contentRetention, transfer, or performance at work
Project descriptionA skill appears in work contextPossible use in a specific projectThe person’s contribution, quality, or autonomy
Manager endorsementA named reviewer supports a claimValidation by that reviewer under some criteriaObjective or portable proficiency
Assessment resultPerformance under a stated instrumentEvidence against that assessment’s rubricPerformance in every real operating context

This is not an argument against inference. Profiles are sparse and titles hide the work inside a role. The OECD’s current skills-first framework places emphasis on demonstrated skills and calls for robust identification and assessment methods. Inference finds possible signals at scale. It cannot provide proof it never collected.

The eight separations a trustworthy system preserves

Many implementations fail in the data model. They place a skill ID beside a person ID, add a score, and call it intelligence. The score travels until nobody remembers what it meant.

A responsible system keeps eight concepts separate:

ConceptQuestion it answers
ObservationWhat event, text, result, or assertion was actually recorded?
InferenceWhat additional claim did a model derive from that observation?
ValidationWho reviewed the claim, against which criteria, and with what result?
ProficiencyWhat can the person demonstrably do, under which rubric and conditions?
ConfidenceHow often is this model expected to be correct for comparable claims?
ProvenanceWhich sources, transformations, taxonomy version, and model produced the claim?
RecencyWhen was the underlying evidence observed, and when should it be reviewed or expire?
Decision rightsWhich workflows may use the claim, and who remains accountable?

Confidence is not proficiency. Validation is not source provenance. Recency is not a cosmetic “last updated” label. And an employee accepting a suggested skill does not retroactively make the model’s original inference an observation.

The practical design rule is to store a skill claim, not a boolean “has skill” field:

subject + relation + skill + evidence + time + method + confidence + validation + permitted use

A record might say employee E-1042 possibly used “SQL query optimisation”; a dated project artefact was mapped to skill S-881 by extractor 3.2; calibrated confidence was 0.81; no expert reviewed it; and it may inform learning but not staffing. That is inspectable. “SQL: 81%” is not.

A defensible inference pipeline

The best implementation does not begin with a model. It begins with a decision and the cost of being wrong.

1. Set the decision boundary

“Recommend optional learning” and “exclude someone from a project” cannot share a threshold. A wrong course suggestion is a nuisance; a wrong staffing claim can create operational risk or deny an employee an opportunity.

Define the action, error costs, owner, and prohibited uses before ingesting data. If the team cannot name the decision, it is building surveillance capacity in search of a purpose.

2. Capture evidence events without changing their meaning

Store the source, subject, timestamp, permissions, and original relationship. A course remains a completion. A CV statement remains a self-declaration. A code contribution does not become an automatic rating of code quality.

Data minimisation matters here. “Available to the model” is not the same as appropriate to use. Private communications or activity traces may be accessible and still be inappropriate for employee profiling.

3. Extract, then normalise

Extraction asks, “what phrase appears?” Normalisation asks, “which concept does it mean?” “R” may be a programming language, product code, or one character in a table. “People management” can be a responsibility or merely a course topic.

Keep raw phrases and candidate mappings so a reviewer can reconstruct the choice. The AI Skills Atlas uses stable concepts and explicit prerequisite relationships for the same reason: structure is useful only while its boundaries remain visible.

4. Generate a typed claim

The model should predict a relationship, not merely attach a skill. Useful relationship types include mentioned, used_in_context, self_declared, required_by_role, adjacent_to, and candidate_for_validation. Interest, exposure, capability, and certification must not collapse into one “skill strength” score.

Require the system to abstain when the text is ambiguous, the taxonomy has no good match, or the evidence is outside its evaluated domain. An unknown is a valid output.

5. Calibrate confidence

A confidence score needs an empirical interpretation: among comparable claims assigned 0.8, roughly 80% should survive review. Modern neural networks can be badly miscalibrated even when their classification accuracy looks strong, as the foundational calibration study by Guo and colleagues demonstrated. Confidence without calibration is interface decoration.

Calibrate by source, skill family, language, and population. A score tested on English technology CVs cannot be assumed to mean the same thing for Polish manufacturing records.

6. Route claims to proportionate validation

Validation should match the decision. Employee confirmation may improve discovery; a domain expert may review a role map; a work sample may support proficiency. Regulated work may require a current external credential or supervised practice.

“Human in the loop” is not a control if that human sees only a score or is rewarded for agreeing quickly. Record the reviewer, evidence, rubric, outcome, and disagreement.

7. Activate only the state the workflow is allowed to use

Learning may consume suggestions. Project search may use employee-confirmed claims while distinguishing them from assessed capability. Consequential workflows should not turn inference alone into a gate.

This separation is technically achievable. Oracle describes Dynamic Skills as identifying, inferring, and recommending skills from organisational data. SAP lets employees add or reject recommended skills, and its 2026 Skills Governance workflow separates imported from inferred skills before publication. These are product descriptions, not proof of accuracy, but they show that explicit states and review gates can be productised.

8. Feed corrections back without corrupting the benchmark

Corrections should propagate and create labelled error events, not rewrite provenance. Keep evaluation data separate from training data, version the model and taxonomy, and retest changes.

Where skills inference fails

The most damaging failures are rarely spectacular hallucinations. They are ordinary category errors repeated at enterprise scale.

Mention becomes mastery

A model finds “Kubernetes” in a profile. The dashboard displays Kubernetes proficiency. The staffing tool ranks the employee for a production migration. Three transformations have occurred, but only the first is supported. This is evidence laundering: each downstream interface removes another qualifier until the hypothesis looks like fact.

Adjacency becomes possession

Skills graphs connect related capabilities, making them useful for discovery and dangerous for attribution. Python may suggest pandas; cloud engineering may suggest infrastructure as code. Neither relationship awards that skill to a person.

A title becomes a stereotype

Role-based inference imports assumptions from job architecture. Titles differ by company, country, seniority, and period. It can also create a circular system: people receive opportunities because their title implied a skill, then those assignments are used as evidence that the title was predictive. The role dictionary deliberately derives role views from skill profiles; it should not be read backwards as proof about every person with the title.

A clean label hides a bad mapping

Normalisation can merge distinct practices or split one into artificial micro-skills. Analytics then appear precise because labels are tidy, not because the mapping is valid.

Historical opportunity becomes “merit”

Work histories reflect who received projects, training, and visibility. A model can reproduce that allocation pattern while appearing skills-based. Evaluate who is missing from the evidence. The US Equal Employment Opportunity Commission warns that algorithmic employment tools can screen out people with disabilities when their inputs or scoring methods do not reflect how qualified people actually perform work; its AI and ADA guidance is a useful procurement reference even outside recruitment.

Time disappears

A skill used once in 2019 and one exercised weekly cannot share a status. Evidence needs a date and domain-specific review rule; a generic “skill half-life” is no substitute for policy.

How to evaluate a skills-inference system

A vendor’s global accuracy figure is almost useless without the unit of prediction, population, threshold, and error distribution. Run a local evaluation before connecting inference to a live workflow. If this evaluation sits inside a procurement, use it alongside the vendor-neutral platform scorecard.

Build a decision-specific reference set

Sample across job families, levels, locations, languages, evidence density, and unusual career paths. Split by time or subject so near-duplicates and the same person’s history do not leak between development and test sets.

Have two qualified reviewers label the same sample against a written rubric; measure agreement and adjudicate differences. A “gold set” built from one reviewer’s intuition is gold-painted. NIST’s AI Risk Management Framework calls for documenting experimental design, data suitability, representativeness, and construct validation, then demonstrating that a system is valid and reliable for its conditions of use in the AI RMF Core.

Report the metrics that expose trade-offs

MeasureWhat it revealsQuestion for the pilot
PrecisionHow many accepted inferences were correct?At the operating threshold, how much false attribution remains?
RecallHow many reference skills were found?Whose capabilities remain invisible?
CalibrationDoes confidence match observed correctness?Is “0.8” meaningful for this source and population?
Coverage / abstentionHow often does the model make no claim?Does apparent completeness come from guessing?
Top-k relevanceAre the first suggestions useful?Can employees review the list without fatigue?
Reviewer agreementIs the target construct consistently defined?Are model errors distinguishable from rubric ambiguity?
Slice performanceWhere do errors concentrate?What changes by language, role, location, or evidence density?
Stability and driftDoes performance survive change?What happens after a taxonomy, model, or work pattern changes?

Publish sample sizes and uncertainty. Inspect errors qualitatively: acceptable average precision can conceal failure in the role family the pilot exists to support.

Test the workflow, not only the model

Run in shadow mode first, without changing access to work. Test whether users understand labels, can find evidence, correct errors, and know who acts on a correction.

Finally, measure the decision outcome: recommendation relevance, staffing quality, time saved, opportunity distribution, or post-learning performance. Profile completion and dashboard visits measure activity, not validity.

Decision rights should rise with evidence

One inference engine can support several workflows, but each needs a different gate:

UseReasonable minimumInference alone?
Catalogue or content taggingTraceable extraction with sampled quality reviewOften acceptable
Optional learning discoveryCalibrated suggestion, explanation, easy rejectionUsually acceptable
Employee profile suggestionSource visibility, employee correction, no silent downstream useAcceptable as a suggestion
Internal opportunity discoveryConfirmed interest or experience plus transparent matchingUseful for inclusion, not exclusion
Project staffingCurrent, role-relevant evidence and accountable reviewInsufficient
Hiring, promotion, pay, termination, safety authorisationValidated evidence, appropriate assessment, human accountability, legal reviewNot defensible as the deciding evidence

These are implementation guardrails, not a substitute for jurisdiction-specific legal advice. In the EU, the classification turns on the system’s intended purpose and use. The consolidated EU AI Act lists specified recruitment and worker-management systems in Annex III, including systems used for certain employment decisions, task allocation based on individual behaviour or traits, and performance monitoring. Following the AI Omnibus, the European Commission says the Annex III high-risk rules apply from 2 December 2027.

That date is not permission to postpone governance. The GDPR already applies to personal data, profiling, and solely automated decisions with legal or similarly significant effects; its Article 22 safeguards include human intervention and the ability to contest qualifying decisions. National employment, equality, data-protection, consultation, and sector rules may add further requirements.

What to demand from an implementation

Before buying or building, require a working demonstration of these controls on your own messy data:

  1. A claim schema that separates observation, inference, validation, proficiency, confidence, provenance, recency, and permitted use.
  2. A source register showing what data is used, why, under which permissions, and for how long.
  3. Stable internal skill IDs, taxonomy versioning, and reversible source-to-concept mappings.
  4. Evaluation results for your population, sources, languages, and actual operating threshold.
  5. Precision, recall, calibration, abstention, slice results, sample sizes, and reviewed errors.
  6. A visible explanation that names the evidence—not a generic paragraph generated after the score.
  7. Employee correction and contest routes, with tested propagation to every downstream system.
  8. Separate policies for learning, discovery, staffing, selection, workforce planning, and any other use.
  9. Named owners for taxonomy, model evaluation, source data, privacy, domain validation, and each consequential decision.
  10. Audit history for source changes, model and taxonomy versions, reviews, overrides, exports, and use in decisions.
  11. Monitoring thresholds, review dates, incident handling, and a kill switch for each workflow.
  12. Contractual access to export claims, evidence references, corrections, and identifiers in a usable form.

If the demonstration jumps from an attractive profile to a business outcome without showing the claim ledger in between, ask to see the missing layers. If the supplier cannot expose them, assume your organisation will have to build that governance around the product.

The real advantage is disciplined uncertainty

AI makes skills discovery cheaper. It does not make capability measurement free. The strongest system is not the one that fills every profile; it is the one that knows which claims are observed, which are inferred, which have earned validation, which have expired, and which must never decide anything on their own.

That design also produces better strategy. The AI Exposure research shows why a single role score needs skill-level context. The same discipline applies inside an organisation: retain the evidence beneath the score, expose uncertainty, and make the decision owner visible.

An inferred profile should be treated as a search index over hypotheses, not a personnel record. Build that distinction into the data model, workflow, and contract before scale makes it costly to recover. If you are designing the claim model, local evaluation, or governance workflow, get in touch to discuss the implementation.

Continue exploring