Why skill taxonomies cannot sequence reskilling: build a learning DAG
A company decides that 120 software engineers need to become capable of building AI-enabled products. Its catalogue contains Python, deep learning, transformer architecture, retrieval-augmented generation, model evaluation, and MLOps. There is a course for every label, but no answer to the question on which the investment depends: who should learn what next?
A flat skills list can describe the destination. It cannot sequence the journey. If everyone is assigned the same curriculum, experienced engineers repeat what they know, novices reach advanced material without the foundations to diagnose failure, and course completion becomes a substitute for capability.
This is the missing layer in many reskilling programmes. A taxonomy gives work a shared language; an ontology adds relationships; a prerequisite graph turns a target capability and a starting point into a testable learning route. It is more useful than a list because it contains direction. It is also more dangerous, because an unjustified edge can make one person's preferred route look like an objective law.
A prerequisite graph is not a map of truth. It is a versioned set of claims about learning order, for a defined outcome, audience, and level of proficiency.
Taxonomy, ontology, and prerequisite graph are not synonyms
The three structures can share identifiers and infrastructure, but they answer different questions.
| Structure | The question it answers | Typical relation | What it does not establish |
|---|---|---|---|
| Skills taxonomy | What concepts do we recognise, and how are they named and grouped? | “Transformer architecture” belongs to “Generative AI”. | Which concept should be learnt first. |
| Skills ontology | How do skills, tasks, roles, tools, evidence, and other concepts relate? | A task uses a skill; a skill is adjacent to another; a role requires a skill. | That one relation implies a learning sequence. |
| Prerequisite graph | What capability usually needs to be in place before a learner can reach a declared target? | Linear algebra precedes transformer architecture for a specified learning outcome. | That the route is necessary for every learner or every use of the target skill. |
This distinction matters in implementation. “Related to” is symmetric and weak: SQL can be related to data engineering without either being a prerequisite for the other. “Part of” describes composition. “Similar to” supports discovery. A prerequisite edge is directed and consequential. It changes recommendations, assessment order, budget, and sometimes access to opportunity.
An edge needs a verb, a strength, and a boundary
Write an edge as a sentence: A should normally be demonstrated before B, for audience C to reach outcome D at level E. If the team cannot complete that sentence, the edge is probably only an association.
The current AI Skills Atlas uses three strength bands:
- Hard: 166 relations (56.8%).
- Medium: 113 relations (38.7%).
- Soft: 13 relations (4.5%).
These labels are editorial judgments about learning order, not measured causal effects. They are most useful as navigational constraints that a learner can challenge when their background differs from the assumed path.
The labels are deliberately qualitative.
- A hard prerequisite means that skipping it usually prevents the declared outcome or makes failure impossible to diagnose. It is a default gate, not a claim that no exceptional route exists.
- A medium prerequisite materially improves comprehension, transfer, or speed, but a learner may compensate through experience, support, or a different route.
- A soft prerequisite is useful context. It should influence a recommendation without blocking progress.
Strength is not confidence. “Hard” describes the proposed dependency; confidence describes how well evidence supports that proposal. A hard edge based on one expert's intuition can have low confidence. A soft edge observed repeatedly across cohorts can have high confidence. Collapsing the two dimensions creates false precision.
Someone calling a managed language-model API does not need the same foundations as someone implementing attention or debugging training. The label may be identical while the required performance—and therefore the route—is different.
Why the learning layer is usually a DAG
A directed acyclic graph, or DAG, gives each edge a direction and contains no closed loop. If linear algebra precedes deep learning, and deep learning precedes transformer architecture, the graph cannot also require transformer architecture before linear algebra.
That constraint is operationally valuable. It allows a system to:
- walk backwards from a target until it reaches capabilities already demonstrated;
- put the remaining nodes into a defensible order;
- identify foundations shared by several target capabilities;
- calculate alternative routes when a learner can bypass or validate a node;
- show where no viable path exists with the current content and assessment supply.
Human learning is not itself acyclic. Practice changes understanding; advanced examples can make a basic concept intelligible; neighbouring capabilities reinforce one another. The DAG is a planning abstraction, not a theory of cognition. A cycle in the data therefore needs interpretation. It may reveal a modelling error, two nodes that should be one capability, an iterative learning loop, or two legitimate routes defined at different proficiency levels. Do not delete the least convenient edge merely to make a topology check pass. Resolve what the cycle means.
Transformer architecture: one target, several honest routes
The atlas currently records these direct prerequisites for Transformer Architecture:
- Deep Learning (hard) — Self-attention, layer normalization, residual connections, softmax — all are DL building blocks assembled in the Transformer
- Linear Algebra (hard) — Q·Kᵀ/√d is a scaled dot product of matrices; multi-head attention is parallel matrix projections — Transformers ARE linear algebra in action
The relation to linear algebra is visible in the scaled dot products and matrix projections used by attention. The relation to deep learning is visible in the wider system of learned parameters, backpropagation, residual connections, normalisation, and optimisation. The primary architecture paper, Attention Is All You Need, does not become easy because a learner can run a library implementation. Deep Learning by Goodfellow, Bengio, and Courville sets out much of the mathematical and neural-network foundation required to reason about what the implementation is doing.
But the edge should still not be interpreted as “nobody may touch a transformer until they have completed two academic courses”. Consider three outcomes:
| Target outcome | Proportionate route |
|---|---|
| Integrate a managed model into an application and evaluate its behaviour | API and software foundations, prompting and evaluation may matter before mathematical depth; the learner can use transformer architecture as context. |
| Fine-tune a transformer and diagnose poor training behaviour | Deep-learning practice, optimisation, model evaluation, and the relevant linear algebra become material prerequisites. |
| Design or research a new attention architecture | Mathematical fluency and deep-learning foundations are hard constraints; reading and reproducing primary research becomes part of the route. |
The graph becomes useful only after the organisation states which of these outcomes it means. A generic “Transformer Architecture — required” field hides that decision.
It also needs a placement mechanism. If an engineer can already explain attention, inspect tensor shapes, and diagnose an unstable training run, the pathway should accept suitable evidence and start later. A prerequisite graph that cannot recognise prior learning is merely a fixed curriculum with better graphics.
What an enterprise can do with the graph
The strongest use cases connect a real capability target to a bounded population.
- Sequence reskilling. Generate the next plausible step from demonstrated capability rather than assigning the same catalogue to everyone.
- Find common foundations. Identify prerequisites shared across several target roles and build them once.
- Place assessments. Test high-impact prerequisites early to avoid missing or repeating foundations.
- Compare build, buy, and borrow. The distance to a target can help distinguish a credible internal pathway from a gap that cannot be closed within the delivery horizon.
- Design cohorts and support. Group learners with similar gaps for instruction, mentoring, or protected practice.
- Manage change. When a target capability changes, trace which pathways, assessments, and resources depend on it.
These are planning and development uses. Path length is not human potential, and a missing prerequisite is not proof that a person cannot perform. Do not turn a shortest-path algorithm into an employment gate. Consequential decisions require direct, job-relevant evidence.
Six ways prerequisite programmes fail
- The universal route. One subject-matter expert's education becomes mandatory for everyone, ignoring prior learning, role context, accessibility, and alternative routes.
- Association masquerading as sequence. Every adjacent concept becomes an edge. The graph is dense, plausible, and useless because almost everything appears to precede everything else.
- False mathematical authority. Edge weights such as 0.83 are published without outcome data, calibration, or a definition. The decimal records confidence theatre, not knowledge.
- Catalogue-first modelling. The team maps thousands of concepts before choosing a target capability, audience, or business horizon. Governance debt arrives before evidence of value.
- A frozen pathway. Technologies, work design, learning resources, and learner populations change while the route remains untouched. The graph preserves yesterday's expert consensus.
- Completion as outcome. The programme measures modules finished and nodes “acquired”, but not whether people can perform the target work or whether time to capability improved.
Learning order is not a property stored inside a skill. It is a claim conditioned on the outcome and learner; the data model must preserve those conditions.
The minimum data contract for a useful edge
A pilot does not need a universal ontology. It does need enough information to make each material dependency inspectable.
| Field | Minimum question to answer |
|---|---|
| Prerequisite and target IDs | Which two versioned concepts does the edge connect? |
| Direction and relation type | What exactly should precede what—and is this truly a prerequisite rather than similarity or composition? |
| Strength | Is the edge hard, medium, or soft for the declared outcome? |
| Rationale | What failure, delay, or loss of understanding does the prerequisite prevent? |
| Audience and target proficiency | For whom, and for what observable performance, does the claim apply? |
| Evidence and confidence | Does support come from expert review, curriculum analysis, learner data, outcome analysis, or a mixture? How strong is that support? |
| Validation or bypass rule | What evidence allows a learner to satisfy or challenge the prerequisite? |
| Owner, version, and review trigger | Who can change the edge, what changed, and when must it be revisited? |
Keep rejected edges and exceptions: they show where the pathway does not fit the work. In this project's cross-validation methodology, model agreement is provenance, while prerequisite strength remains editorial judgement. Agreement can nominate an edge for review; it cannot validate the sequence.
A bounded pilot in six steps
Choose one target capability linked to a funded workforce decision—not “build an enterprise skills graph”. Then:
- define the target as observable work, the required level, the population, and the time horizon;
- collect plausible direct prerequisites from domain experts, current training, work artefacts, and primary technical sources;
- classify each edge, record its scope and rationale, and test the subgraph for cycles, duplicate concepts, dead ends, and unreachable targets;
- assess the starting capabilities of a small cohort with proportionate evidence, allowing people to challenge or bypass recommendations;
- run the pathway alongside a reasonable baseline, such as the current standard curriculum; and
- compare outcomes, inspect exceptions and subgroup differences, revise the graph, and publish the next version.
Start with tens of meaningful edges, not tens of thousands of decorative ones. The first purpose of the pilot is to discover where the proposed sequence is wrong.
Metrics that test the graph rather than decorate it
Cycles, orphaned nodes, duplicate paths, excessive depth, and target reachability can expose data defects. They do not prove that a pathway works. Quality needs outcome measures too.
| Measure | What it can reveal |
|---|---|
| Edge challenge and override rate | Where experts or learners repeatedly reject the default route. |
| Diagnostic bypass rate | Whether the graph recognises prior capability instead of prescribing redundant learning. |
| Backtracking or remediation rate | Where a prerequisite is missing, too weak, or placed too late. |
| Time to demonstrated target capability | Whether sequencing improves the result that matters, not merely completion speed. |
| Target assessment success versus baseline | Whether the graph adds value beyond the previous curriculum or recommendation method. |
| Outcome and error distribution | Whether the route works differently by role, location, language, accessibility need, or another relevant group. |
| Edge freshness and review coverage | Whether consequential claims still have active owners and current support. |
No metric crowns the graph “accurate”. A low override rate may mean a good route—or that learners cannot safely object. Read the measures together with qualitative review and demonstrated work.
The graph is a decision layer, not ground truth
A skill list says what an organisation has named. A prerequisite graph proposes how capability can be built. It can reduce wasted training, make reskilling plans more realistic, and expose shared foundations. It also concentrates editorial power in edges most users will never inspect.
The standard is not graph size. It is whether learners and decision owners can understand a route, challenge its assumptions, prove a different starting point, and test whether it produced the target capability.
For a manual starting point, choose one target in the AI Skills Atlas, walk backwards through its hard prerequisites, and stop at the first capability you can demonstrate. Then ask what evidence would justify the next edge for your role and intended outcome. That is already more honest than a generic roadmap.
If you need to turn a capability target into a governed learning DAG, evidence contract, and measurable reskilling pilot, get in touch to discuss the implementation.
Sources and limitations
- Vaswani et al., Attention Is All You Need, for the Transformer architecture.
- Goodfellow, Bengio, and Courville, Deep Learning, for mathematical and neural-network foundations.
- OECD, Building a common skills language, for taxonomy interoperability and governance challenges.
- How the atlas was assembled, including why prerequisite relations and strengths remain editorial claims.
The graph describes plausible learning dependencies in a selective technical atlas. It has not been validated as a causal model of learning, a universal curriculum, or a basis for employment decisions. Strength labels are editorial judgements and must be re-evaluated for the target population and outcome.
Continue exploring
- Browse prerequisites in the AI Skills Atlas
- Understand the decision discipline behind skills intelligence
- Assess whether a programme has earned the right to activate recommendations
- Evaluate skills intelligence platforms with a vendor-neutral scorecard
- Build a decision-grade skills intelligence dashboard
- Read the cross-validation methodology