What three AI research runs can—and cannot—validate
The current edition contains 426 skills across 15 domains, connected by 591 ontology relations and 292 prerequisite relations. The original research runs surfaced 115 skills in all three systems, 78 in two, 122 in one, and 111 in none. Those gaps remain visible rather than being relabelled as consensus.
The first edition of the GenAI 2026 Skills Atlas began with three deep-research runs: Anthropic Opus, OpenAI Deep Research, and Google Deep Think. Each system worked from the research brief without seeing the other two outputs. We then compared what they surfaced, normalised overlapping labels, and recorded the result as a 0/3–3/3 research-run signal.
That design needs a precise claim. Agreement between several models is a signal of discovery stability and provenance; it is not proof that a skill is important, correctly defined, or even factually described. Conversely, disagreement is not a defect to be averaged away. It is often the most useful part of the result because it shows where scope, language, or judgement is unstable.
This research note documents what we did, what the marker means, and how an enterprise team can use a similar method without mistaking an AI ensemble for expert validation.
One limitation should be explicit at the outset. This public edition preserves the provider-family flags and the current dataset, but it does not publish the exact product and model versions, run dates, original research brief, raw outputs, or complete normalisation log. It is therefore a methodological account and provenance note, not a fully reproducible experiment. A future edition should release those artefacts — or publish hashes and timestamps where licences prevent release.
What we actually did
The three systems were used as separate discovery passes, not as a jury issuing a verdict. For each candidate skill, the requested output covered:
- a definition: what the skill is and what problem it solves;
- categorical placement in the broader taxonomy;
- prerequisites: what a learner should know first;
- one recommended book, paper, or course;
- notes on relevance to GenAI and data science in 2026.
We brought the outputs into a comparison spreadsheet, reconciled labels that appeared to describe the same capability, and stored the resulting presence flags with each atlas skill.
Later editorial additions retain 0/3 coverage, while reviewed sources on each skill page form a separate evidence layer. The marker records where an item entered the process; it does not say that three models proved its current description.
Why use separated runs at all?
A single model can produce a coherent list while omitting an entire school of thought. Its answer is sensitive to prompt wording, training material, and how the product retrieves or ranks sources.
Three provider families give us three opportunities to notice a candidate and make omissions visible. If two systems surface causal inference and one does not, the gap becomes inspectable.
“Separated” does not mean statistically independent. The systems may have learned from overlapping web pages, papers, books, benchmarks, and code, so they may repeat the same popular framing or error. Products also change over time. The method reduces dependence on one output; it does not create three independent measurements of reality.
The protocol, step by step
The useful design principle is an audit trail from raw suggestion to published node. The current public release retains the final provider flags, but not every artefact described below. A team repeating the method should preserve the complete trail from the start.
- Fix the question and output fields. Define the domain, time horizon, granularity, and required fields before comparing answers.
- Run each system without the other outputs. Do not ask the second system to critique the first during discovery. That would anchor later answers and make apparent agreement cheaper.
- Preserve the raw wording. “AI red teaming” may mean security testing, governance exercises, model-behaviour evaluation, or all three. Early rewriting destroys that difference.
- Normalise entities, not opinions. Merge spelling variants and genuine synonyms, but do not treat two differently scoped concepts as agreement merely because their names resemble one another.
- Build a presence matrix. For every normalised skill, record whether each run surfaced it. The sum creates the 0/3–3/3 marker.
- Separate integration from validation. Editors can add skills, refine definitions, and attach sources without changing what happened in the original runs.
- Publish the lineage. Keep the marker, sources, editorial notes, and edition date distinct.
Its value lies in preserving boundaries that fluent answers easily blur.
How to interpret 3/3, 2/3, 1/3, and 0/3
Agreement is uneven: 115 of 426 skills have 3/3 research-run coverage, 78 have 2/3, 122 have 1/3, and 111 were added during later editorial integration with 0/3 coverage in the original runs. Consensus is therefore a provenance signal, not a quality score.
| Marker | What it supports | What it does not support | Sensible next action |
|---|---|---|---|
| 3/3 | The skill was stable enough to be surfaced by all three runs. | Correctness, business importance, or complete agreement on scope. | Compare definitions and verify claims against current sources. |
| 2/3 | The candidate was reasonably stable within this research setup. | That the absent system “failed” or the majority is right. | Inspect naming, granularity, and the missing run. |
| 1/3 | One run produced a discovery lead worth considering. | That the skill is fringe, new, or wrong. | Seek independent evidence and check whether it duplicates another node. |
| 0/3 | No positive signal exists in the original three-run matrix; the item was added later. | Rejection, low quality, or low importance. | Judge the editorial case and reviewed sources on their own merits. |
The scale is not a quality score. A 3/3 acceptance threshold would reward familiar concepts and could suppress emerging, specialist, non-English, or organisation-specific capabilities.
Disagreement is data
Three examples from the original runs show why the binary flags are only the start of analysis:
- Bayesian Statistics appeared in the Opus and Google runs but not the first OpenAI run. That absence says little by itself about the field's importance. It tells us that discovery was sensitive to the research path taken by one system.
- AI Red Teaming appeared in all three. Our editorial reading of the retained material was that the scope differed: one framing emphasised adversarial security, another policy and governance, while another combined them. That qualitative comparison cannot be independently reconstructed from the public presence flags alone. It illustrates how a 3/3 marker can coexist with a substantive taxonomy dispute; it should not be read as a public benchmark of the three systems.
- Causal Inference appeared in the Opus and Google runs but not the OpenAI run. Its 2/3 marker creates a review question; it does not settle the answer.
Disagreement can reveal an omission, competing names, one name covering several capabilities, or a genuine difference in judgement. Each needs a different response; a majority vote cannot tell them apart.
Failure modes this method does not solve
Cross-model comparison is useful precisely when its limits remain explicit:
- Correlated sources. Three systems can repeat the same error because their training or retrieval material overlaps.
- Prompt sensitivity. Small changes in framing, list length, or detail can change the result.
- False synonymy. Editorial normalisation can manufacture agreement by merging concepts that practitioners would keep separate.
- Granularity mismatch. One model may name a discipline while another lists its techniques.
- Temporal drift. Closed models and search layers change. Re-running an identical prompt later may not reproduce the original output.
- Plausible fabrication. Agreement cannot authenticate a citation, establish proficiency, or demonstrate a workforce capability.
- Popularity bias. Widely discussed skills are easier to recover than tacit, local, or operational knowledge.
For these reasons, the atlas keeps reviewed sources separate from the AI consensus marker. The approach also aligns with the broader risk-management principle in the NIST AI Risk Management Framework: claims about an AI system should be governed and measured in the context in which a decision will be made.
Reproducibility and corrections
Commercial systems change and their retrieval is opaque, so perfect reproduction is impossible. At the publication layer, however, we can version the dataset, retain provider flags, state the edition and method, and never convert later additions into historical agreement.
The present release does this only partially: readers can inspect the resulting flags and the current editorial dataset, but not replay the original runs or audit every normalisation decision. That distinction matters. Repeatable editorial lineage is not the same thing as experimental reproducibility.
A correction should also identify which layer changed:
- a research-run correction fixes transcription or entity matching in the original matrix;
- an editorial correction changes a definition, category, prerequisite, or interpretation;
- a source correction replaces or qualifies evidence without rewriting the historical marker;
- a new edition may run a fresh protocol and should be reported as a new measurement, not blended invisibly into the old one.
The current atlas follows the essential rule: later additions remain visible as 0/3 rather than inheriting consensus they did not earn. If you find a missing, incorrect, or poorly framed item, please send a correction. A useful correction names the skill, disputed field, and strongest available primary source.
What this means for enterprise skills intelligence
The same pattern applies when models infer skills from CVs, job histories, project records, or work products. Agreement can improve triage, but it must not become evidence by multiplication.
A defensible operating model separates three layers:
- Discovery: models propose that a skill claim may exist.
- Evidence: assessments, observed work, credentials, manager reviews, or verified project outcomes support or challenge that claim, with recency and provenance attached.
- Decision: an accountable person applies a threshold appropriate to the consequence—search, learning recommendation, staffing, promotion, or regulated access.
Agreement is most valuable for routing work. A 3/3 inference might enter a lower-priority review queue; a disagreement might trigger clarification; a high-consequence claim might require observed performance regardless of model consensus. The decision rule should become stricter as the cost of being wrong rises.
That is the practical lesson of the atlas experiment: use multiple models to expose uncertainty, not to disguise it. The output becomes more trustworthy when provenance, dissent, and editorial judgement remain visible as different things.
Continue exploring
- Browse the current AI Skills Atlas and inspect the 0/3–3/3 marker on each skill.
- Compare the entries for Bayesian Statistics, AI Red Teaming, and Causal Inference.
- Read about the project's scope and correction policy.
- See how evidence should shape a skills-intelligence platform.
- Thoughtworks Technology Radar remains one inspiration for the edition-based publication format.
— Skills Intelligence