The skills intelligence maturity model: six stages from spreadsheet to decision system
An organisation launches a skills platform, imports a taxonomy, and persuades most employees to complete their profiles. The implementation is live, the dashboard is full, and the programme is described as mature. Then a business leader asks a practical question: Can we staff the next product launch internally without putting delivery at risk?
The system produces a list of people. Nobody can explain which entries came from demonstrated work, which were inferred from job titles, how current they are, or whether the people are available. The programme has accumulated data without earning the right to act on it.
This is why conventional maturity measures are so misleading. Taxonomy size, profile completion, integration count, feature adoption, and platform deployment all describe activity or infrastructure. They do not establish that an organisation can make a better workforce decision.
Skills intelligence maturity is the class of decision an organisation can safely improve, repeatedly, with evidence it can inspect.
A deployed platform is an asset. A maturity level is an earned permission.
The six-stage model below is an original enterprise diagnostic, not an industry standard or a vendor certification. It starts where most programmes start—with an inventory—and ends with a system that learns from the consequences of its own recommendations. Each stage has a gate. If the gate is not passed, adding features from a later stage creates exposure, not maturity.
Maturity belongs to a decision, not to the enterprise
Before applying the model, choose one decision and one population. An organisation can be mature in recommending optional learning and immature in selecting people for promotion. It can have decision-grade evidence for licensed engineers at one plant and little more than self-reported profiles for the rest of the workforce.
This distinction prevents the maturity assessment from becoming another enterprise-wide vanity score. The unit of analysis should sound like this:
For this population, can we use this evidence to improve this decision, at this level of consequence, and learn whether the intervention worked?
Score the highest stage for which every earlier gate is also operating. Do not average across stages.
| Stage | Organisational capability | Decision the organisation has earned the right to make |
|---|---|---|
| 1. Inventory | Locate bounded claims and their sources | Decide what needs investigation |
| 2. Mapping | Connect claims to shared, governed concepts of work | Decide which records are meaningfully comparable |
| 3. Validation | Establish whether evidence is adequate for a declared use | Decide which claims may enter a bounded analysis |
| 4. Intelligence | Compare supply, demand, risk, and scenarios with visible uncertainty | Recommend an intervention |
| 5. Activation | Put the recommendation into an accountable workflow and measure the result | Act on the intervention |
| 6. Learning system | Use outcomes and corrections to improve the evidence, model, and workflow | Adapt the decision system itself |
A carefully governed spreadsheet pilot can reach stage four for a narrow decision. A global platform can remain at stage one.
Stage 1 — Inventory: we can locate the claims
At the inventory stage, the organisation can find relevant records and say where they came from. That sounds modest because it is. Yet it is already more useful than a profile-completion campaign that merges self-report, course history, inferred skills, credentials, and work experience into one undifferentiated field.
The inventory must have a declared boundary: which population, time period, systems, languages, and evidence types it covers. Missing data is represented as missing, not converted into “does not have the skill”. Source records retain their dates, permissions, and original meaning. A course completion remains a learning event; it does not become proof of applied capability.
Gate to stage 2: for the chosen decision, the team can trace material claims to dated source records, describe coverage and missingness, and identify who is permitted to use the data.
Failure mode — completion theatre. The programme optimises the percentage of people with a profile. Employees learn to enter broad labels, managers chase completion, and the dashboard fills up. Coverage rises while evidential quality remains unknown. The database is more complete; the decision is no safer.
At this stage the correct output is not “we have 600 cloud engineers”. It is “we have 600 records that mention cloud-related concepts, drawn from these sources, with these limits”.
Stage 2 — Mapping: we can compare the claims
Mapping creates a shared language among evidence, tasks, skills, roles, jobs, and business capabilities. The goal is not to force the entire organisation into one perfect taxonomy. It is to make the concepts needed for one decision sufficiently explicit and interoperable.
A useful mapping layer has stable identifiers, definitions, aliases, version history, and named owners. It can preserve a local plant term or professional standard while mapping it to a broader enterprise or external concept. It supports “unmapped” and “ambiguous” as legitimate states. When a concept is split, merged, or retired, the organisation can inspect what changed downstream.
This is consistent with the OECD's finding that a common skills language is foundational while granularity, duplication, interoperability, and maintenance remain substantial design challenges. Its 2026 analysis also warns that occupational classifications can hide variation within jobs and delay recognition of emerging skills. A shared vocabulary is therefore infrastructure for better questions, not a substitute for evidence. See the OECD's chapter on building a common skills language.
Gate to stage 3: domain owners agree that the mapped concepts distinguish the work that matters; changes are versioned; ambiguous records remain visible; and role or task mappings have a review route.
Failure mode — the ontology cathedral. A central team spends years modelling every conceivable skill before testing whether the language improves a decision. The catalogue becomes too costly to challenge and too generic to represent local work. More concepts create more governance debt, not more intelligence.
Stage 3 — Validation: we know what the evidence permits
Validation is the point at which a skills programme stops treating all claims as equivalent. It separates observation, inference, validation, proficiency, confidence, and recency, then applies an evidence threshold proportionate to the decision.
An employee-confirmed suggestion may be adequate for career exploration. It is unlikely to be adequate for appointing a technical authority, determining pay, or assigning safety-critical work. The same record can therefore be permitted in one workflow and prohibited in another.
If AI inference is used, validation requires a representative evaluation set, predefined labels, error analysis at the operating threshold, and a route to abstain. If human review is used, the organisation records who reviewed what, against which rubric, with which evidence. A manager click is not automatically ground truth. Employees and experts need practical correction and challenge paths, and corrections must propagate to downstream uses.
The NIST AI Risk Management Framework Core is useful here because it ties measurement and governance to context, intended purpose, risk tolerance, accountability, and continuing review. It does not invite organisations to declare a model “valid” in the abstract. Neither should a skills programme.
Gate to stage 4: the decision has an explicit evidence policy; material claims can pass, fail, or remain unresolved against it; model and human errors are reviewed; and affected people have a usable correction or contest route.
Failure mode — validation theatre. The organisation adds a badge called “verified” without preserving the method, rubric, scope, or date. Inference is accepted by an employee and quietly reclassified as observation. A confidence score is displayed as proficiency. Validation becomes a cosmetic state rather than a defensible process.
Stage 4 — Intelligence: we can recommend an intervention
At stage four, the programme connects validated, bounded evidence to a decision. Supply and demand are expressed in compatible units and time horizons. Availability, aspiration, location, prerequisites, and substitutability are modelled where they materially affect the choice. Missing evidence and contradictory signals remain visible.
The output is comparative rather than declarative. Instead of announcing a “skills gap”, the team can compare scenarios: develop adjacent internal capability, change the work design, recruit, contract, defer, or accept the risk. Assumptions and sensitivity are shown. The recommendation has a baseline and an owner who understands both the business context and the evidence boundary.
This is where a dashboard can become intelligence, but only if it makes the reasoning inspectable. If a leader cannot tell which assumption changed a recommendation, the visualisation has compressed the problem rather than explained it.
Gate to stage 5: a named decision owner can use the analysis to choose among real interventions; the recommendation exposes its assumptions, uncertainty, and evidence boundary; and domain experts agree that the alternatives reflect operational reality.
Failure mode — insight without consequence. The analytics team produces gap heatmaps, adjacency networks, and predictive scores that never alter a budget, workforce plan, staffing decision, or learning investment. The organisation has built an observatory, not a decision capability.
Stage 5 — Activation: the recommendation changes a real workflow
Activation connects intelligence to the place where work is allocated and resources move. A recommendation can trigger a validation, open an internal opportunity, alter a hiring requisition, fund a learning pathway, change a team design, or inform a build-buy-borrow decision. The workflow has an accountable human owner, documented overrides, and a stop condition.
Crucially, the organisation measures the result against a baseline. It distinguishes activity from outcome: a match is not a move; a course is not capability; a profile view is not a better staffing decision. It also examines the distribution of opportunity and error, not only the average result.
Higher consequence requires a harder gate. The EU AI Act identifies specified AI uses in recruitment, selection, promotion, termination, task allocation based on personal characteristics, and worker monitoring or evaluation as high-risk, subject to the Regulation's scope and conditions. The official Regulation is a reminder that risk follows intended use, not the product label. Legal classification requires case-specific advice; the operational principle is simpler: do not let a low-evidence discovery feature become an employment gate by accident.
Gate to stage 6: the recommendation operates in a real workflow; decision rights, overrides, and prohibited uses are enforced; business and worker outcomes are measured against a baseline; and the organisation can stop or reverse the intervention.
Failure mode — the recommendation graveyard. The platform produces plausible matches, but managers cannot release people, budgets do not move, opportunities are fictional, or employees do not trust the process. The model may work as designed while the operating model guarantees that nothing happens.
The more dangerous variant is premature activation: an unvalidated score enters selection or task allocation because integration made it easy. The worst maturity gap is not missing skills data; it is an intervention that has outrun its evidence.
Stage 6 — Learning system: outcomes change the system
A stage-six capability learns from what happened after the decision. It captures whether the person moved, whether capability was demonstrated, whether delivery improved, which recommendation was rejected, which claim was corrected, and which constraint turned out to matter. Those events become governed feedback, not anecdotes lost in project retrospectives.
The organisation can distinguish several possible failures. Perhaps the evidence was poor. Perhaps the mapping represented the work badly. Perhaps the recommendation was sound but incentives blocked the action. Perhaps the action occurred but did not address the real business constraint. Each diagnosis implies a different change.
Models, taxonomies, evidence policies, thresholds, and workflows are versioned and re-evaluated. Performance is monitored as the population and context change. Retired concepts and rejected inferences remain auditable. Resources for governance and correction are treated as operating cost, not temporary implementation work.
This reflects Cedefop's definition of skills intelligence as a process of identifying, collecting, analysing, synthesising, and presenting information—and its explicit qualification that the process must be continuous and iterative to remain relevant. A static inventory can document yesterday accurately and still recommend the wrong action tomorrow.
Stage-six gate: evidence, decisions, interventions, corrections, and outcomes form a traceable loop; material changes trigger re-evaluation; and the organisation can show where feedback changed the system rather than merely adding more data.
Failure mode — the self-fulfilling model. The system learns only from people who received an opportunity. It then treats their subsequent visibility as proof that its original ranking was correct. Historical access becomes future “merit”. A learning system needs counterfactual thinking, rejected and missed cases, independent review, and deliberate exploration—or it will perfect its own blind spots.
The 18-question maturity diagnostic
Apply these questions to one decision. A “partly” is a “no” until there is inspectable evidence. The first stage with a failed gate identifies the immediate constraint; the answer is not to total the ticks and calculate a percentage.
| Stage | Three gate questions |
|---|---|
| 1. Inventory | Can we trace every material claim to a dated source? Can we describe the population and missingness? Are permissions and permitted uses known? |
| 2. Mapping | Are decision-critical concepts defined and versioned? Have domain owners reviewed the relationships? Can the model retain ambiguity and local meaning? |
| 3. Validation | Is there a decision-specific evidence threshold? Are inference, confidence, and proficiency kept separate? Can affected people inspect and challenge material claims? |
| 4. Intelligence | Is supply compared with demand in compatible units and time? Are assumptions and uncertainty visible? Does the analysis distinguish real interventions? |
| 5. Activation | Is a named owner authorised to act? Are baseline, outcome, override, and stop conditions defined? Can we detect who benefits, who is missed, and who bears error? |
| 6. Learning system | Do outcomes and corrections update the model? Are changes versioned and re-evaluated? Can we separate evidence failure from workflow or incentive failure? |
Three patterns deserve particular attention:
- Wide but shallow: enterprise coverage is high, but the first failed gate is validation. Narrow the use case before collecting more profiles.
- Technically advanced but operationally inert: analytics pass stage four, but no owner can move people, work, or budget. Fix decision rights and incentives before buying another feature.
- Activated but ungoverned: recommendations affect people despite weak evidence or no challenge route. Restrict the workflow immediately; this is maturity debt with human consequences.
A thought experiment: staffing a regulated claims-automation programme
Consider a hypothetical insurer deciding how to staff a claims-automation programme. This is not a case study and claims no measured result; it shows how the decision changes as each gate is passed.
At inventory, the team locates project histories, credentials, role records, assessments, and employee-declared interests. It does not count every mention of “automation” as capability.
At mapping, broad labels are decomposed into the work the programme needs: claims operations, process design, data engineering, model evaluation, fraud controls, privacy, customer escalation, and vendor management. Local rules and systems are mapped without pretending that a generic AI taxonomy contains the whole operating context.
At validation, evidence rules vary by use. A declared interest can open a development conversation. Recent project evidence and expert review may support staffing consideration. Specified responsibilities may require current credentials or demonstrated performance. Inferred skills remain hypotheses until the relevant gate is passed.
At intelligence, the organisation compares internal deployment, focused validation, reskilling, external recruitment, contracting, work redesign, and sequencing. It can see where the apparent gap comes from missing evidence, scarce capability, unavailable people, or unrealistic demand.
At activation, an accountable portfolio owner changes the staffing and development plan. The programme records why candidates were considered, where human judgement overrode the model, and which operational and worker outcomes will test the decision.
At learning-system maturity, delivery evidence, assessments, corrections, moves, and failures update the mappings and thresholds. If the real bottleneck was claims authority or manager release rather than technical skill, the next recommendation reflects that finding.
The organisation has not become “stage six” for every purpose. It has built a stage-six loop for one bounded decision. Promotion, redundancy, or plant safety would require their own evidence, controls, and gates. That is a strength of the model, not a limitation: maturity should narrow a claim to what has been demonstrated.
How to move up one stage without starting a transformation programme
Choose the smallest decision whose outcome matters enough for a business owner to care and whose consequences are bounded enough to learn safely. Then invest in the next missing gate:
- From inventory to mapping, model only the concepts required by that decision.
- From mapping to validation, create a representative reference set and evidence policy before adding more inference.
- From validation to intelligence, compare real interventions rather than publishing another gap count.
- From intelligence to activation, place the output inside one owned workflow with a baseline and stop condition.
- From activation to learning, capture corrections and outcomes in a form that can change the next version.
Do not purchase scale before identifying the control being scaled. Do not automate a decision the organisation cannot yet explain manually. And do not label the entire enterprise with one maturity number when the evidence supports different permissions for different uses.
The practical ambition is not to reach stage six as quickly as possible. It is to avoid operating above the level the evidence has earned. A mature organisation can say “we do not know”, restrict a claim to a safer use, and design the next test. That restraint is not a lack of intelligence. It is what makes the intelligence credible.
If you want to assess a real workflow, identify its next gate, or turn a broad skills programme into a bounded implementation roadmap, get in touch to discuss the work.