← Back to blog
Published

How to evaluate a skills intelligence platform

by Skills Intelligence

The most dangerous screen in a skills-platform demo is often the one that looks complete: every employee, every skill, every gap, all rendered with decimal-point confidence. Some of that picture may be observed. Some may be imported, self-reported, inferred, stale, or simply inherited from a generic job profile. When the interface erases those distinctions, the precision of the display has exceeded the precision of the evidence.

That is why the first procurement question should not be, “Which platform has the largest skills catalogue?” It should be:

Which decision will this system improve, and what evidence would let us trust that improvement?

A skills intelligence platform should turn fragmented evidence about work and capability into a better hiring, learning, mobility, staffing, or workforce-planning decision. It should not merely rename a catalogue, add an AI search box, or place opaque scores in a dashboard. The catalogue is a component. An inferred score is a hypothesis. The product you are buying is an evidence and decision system.

This is an independent evaluation guide, not a vendor ranking. Skills Intelligence does not sell a platform. The framework below is designed to make a selection and pilot falsifiable: a product should be able to show what it knows, how it knows it, where it is uncertain, and whether its output improves a real workflow. If the category is new to you, start with What is skills intelligence?.

Platform, dashboard, and system of record are not synonyms

A dashboard presents measures. A system of record stores governed information. A skills intelligence platform may include both, but its distinctive job is to connect evidence, vocabulary, relationships, inference, validation, and action.

This distinction prevents a common buying error. A beautiful skills map can rest on stale labels, weak matching, or no evidence of proficiency. Conversely, a modest interface can support a valuable decision if the concepts, evidence, and validation are sound.

The market makes this harder because “skills intelligence” is not one product category. Vendors approach it from different starting points:

Starting pointTypical strengthRepresentative examplesThe question buyers often miss
HCM-embedded skills layerIdentity, job architecture, permissions, and native HR workflowsWorkday Skills Cloud, SAP Talent Intelligence Hub, Oracle Dynamic SkillsIs the native model good enough for our decision, and can we use or export it outside the suite?
Skills data and inferenceNormalisation, extraction, external market signals, and profile enrichmentTechWolf, LightcastWhat is actually observed, and what is predicted from titles, text, or neighbouring concepts?
Talent intelligence and marketplaceMatching people to roles, projects, gigs, mentors, and career pathsEightfold, GloatDoes a good match produce a real move, or merely another recommendation?
Learning-led skills layerConnecting gaps to content, pathways, academies, and developmentDegreed Skills+ and learning-suite alternativesDoes completion create evidence of learning, or is it being treated as evidence of proficiency?
Assessment-led evidenceTests, simulations, work samples, and structured proficiency signalsWorkera, iMocha, HackerRank, and specialist assessorsIs the measure valid for this role, language, population, and consequence?

These are centres of gravity, not permanent boxes. Product roadmaps overlap and acquisitions can change packaging faster than an enterprise procurement cycle. Do not compare every vendor against the same feature spreadsheet. First decide which layer you need to own and which problem you are actually paying a vendor to solve.

Start with a decision contract, not a feature list

Before the RFP, write a one-page decision contract. Name the population, the workflow, the decision owner, the evidence allowed, the people affected, the potential harm, the baseline, and the outcome you want to change. “Create a skills inventory for 40,000 employees” is not a decision. “Reduce time to staff cloud-migration projects without lowering delivery quality or excluding non-obvious internal candidates” is.

The contract also defines what the system must not do. A learning pilot may permit broad, low-confidence recommendations. Promotion, candidate filtering, task allocation, or redundancy decisions need a much stronger evidence threshold and legal review. The same algorithm can be a helpful discovery tool in one workflow and an unacceptable decision proxy in another.

The EVIDENCE framework for platform evaluation

Use eight dimensions to evaluate the system beneath the demo. Together they spell EVIDENCE.

E — Explicit decision and consequence

Require the vendor to configure one real decision, not a generic journey. Identify who sees the output, who can act on it, and how consequential the action is. A platform that performs well for course discovery has not thereby proved that it is valid for staffing or selection.

Test: ask the vendor to state, in one sentence, what its score permits a manager to conclude—and what it does not.

V — Verifiable evidence and provenance

Inspect which sources the platform can use: job architecture, work history, projects, assessments, learning records, CVs, credentials, or external occupational data. More sources are not automatically better. Each assertion should retain its source, date, permission, and method. A CV mention does not establish proficiency; course completion does not demonstrate transfer to work.

Test: select ten displayed skills and trace each one to its contributing records without giving the analyst unrestricted access to personal data.

I — Interoperable language

The data model should support stable identifiers, definitions, aliases, hierarchy, relationships, and version history. It should preserve local meaning while mapping to external references where useful. ESCO, for example, is a free multilingual classification of occupations and skills, but it is a reference language—not a replacement for plant-specific, professional, or organisation-specific concepts.

Test: rename a skill, split it into two concepts, merge a duplicate, and retire another. Then inspect what happens to profiles, role requirements, history, reports, and integrations.

D — Demonstrated model validity

Treat inference and matching as models, not magic. Ask what evidence produced each output, which population was used for evaluation, and how performance changes at the threshold your workflow will use. “Our model is 90% accurate” is meaningless without the task, denominator, labelled set, error costs, and operating threshold. Request precision and recall, confidence calibration, subgroup results, and examples of false positives and false negatives.

For a detailed claim model and evaluation pipeline, see AI skills intelligence: why inference is not evidence.

Test: create a blinded, representative sample from your organisation and have domain experts label it before the vendor sees the answers. Do not accept a benchmark selected after the results.

E — Employee agency and expert validation

People should be able to inspect inferred data, understand its source, correct it, and contest its use. Domain experts need a governed route to challenge taxonomy and role mappings. A “human in the loop” is not meaningful if the reviewer lacks time, authority, context, or a usable explanation.

Test: have an employee correct an inference and follow the correction through every downstream recommendation, dashboard, and export. Record where the original claim remains.

N — Necessary governance and security

Skills data can affect access to work. Review role-based access, purpose limitation, retention, audit history, model updates, incident handling, accessibility, vendor subprocessors, and regional requirements as product capabilities—not as a security appendix. The NIST AI Risk Management Framework offers a useful govern–map–measure–manage structure for this work.

Test: identify every consequential workflow consuming inferred skills and disable one without destroying the governed source records.

C — Connected action and measured outcome

The platform should improve the selected workflow: learning allocation, project staffing, role redesign, internal mobility, or workforce planning. Inspect who acts, what is written back to operational systems, and whether managers must maintain a second profile. Logins, profiles completed, matches generated, and dashboard views are adoption signals—not business outcomes.

Our companion guide defines 12 metrics for a decision-grade skills intelligence dashboard, separating evidence quality, capability risk, workflow activity, and business results.

Test: compare the new workflow with a baseline on time, quality, conversion, and distribution of outcomes. Measure whether a recommendation became an opportunity and whether the opportunity produced value.

E — Economics and exit

Price the operating system, not only the licence. Include integration, taxonomy maintenance, validation, assessments, change management, privacy review, employee support, model monitoring, and internal governance. A proof of concept that relies on vendor consultants or manual data cleanup should cost that work into production. Also test export: concepts, identifiers, relationships, evidence references, confidence, history, and user corrections should leave in a usable form.

Test: ask both teams to rehearse contract exit. If the answer is a CSV of skill names and scores, you do not have portability; you have a souvenir.

A pilot scorecard that attractive features cannot game

Score evidence from your own data and users on a 0–3 scale:

  • 0 — asserted: described in marketing or a sales response;
  • 1 — demonstrated: shown in a controlled vendor environment;
  • 2 — evidenced: reproduced with your representative data;
  • 3 — operational: reproduced in a production-like workflow with documented limits.
DimensionWeightEvidence required for a credible pilot
Explicit decision10Named owner, baseline, consequence, prohibited uses, and success measure
Verifiable evidence15Sample outputs resolve to dated and permitted source records
Interoperability10Versioning and round-trip export work with stable identifiers
Demonstrated validity20Predefined evaluation set, threshold metrics, error review, and subgroup analysis
Employee agency10Explanation, correction, contest, and propagation tested end to end
Necessary governance15Access, retention, oversight, update, incident, and audit controls exercised
Connected outcome15Workflow improves against the baseline without a material distributional harm
Economics and exit5Production TCO and a viable migration path are documented

Multiply each 0–3 score by its weight, then divide the total by three for a percentage. The number is a decision aid, not a substitute for judgment. Set hard gates before the pilot: legal permissibility, traceability, user correction, and acceptable model validity should not be traded away because the interface is excellent. Publish the threshold and the evidence owner in advance.

A responsible 90-day pilot

Weeks 1–2: frame the decision

Choose one population and workflow. Record the baseline, potential harm, accountable owner, data boundary, and the result that would cause you to stop. Form the evaluation set now, before anyone can tune the test to the product.

Weeks 3–5: test the information layer

Map a representative sample. Have domain experts review concepts and inferences. Measure errors at the operating threshold, not at a convenient generic threshold. Include sparse, unusual, multilingual, non-linear-career, and accessibility cases.

Weeks 6–9: run the workflow

Give real users explanations and correction paths. Compare the new process with the baseline on quality, time, conversion, and distribution of outcomes. Keep a failure log. A successful pilot that hides exceptions is only a successful demo.

Weeks 10–12: decide and document

Review subgroup results, governance workload, interoperability, total operating cost, and failure modes. Document where inference may and may not be used. If the pilot continues, set re-evaluation dates: jobs, evidence, taxonomies, integrations, and models all change.

Questions to put into the RFP

Do not ask only whether a feature exists. Ask the vendor to supply an artefact or run a test.

  1. What exactly is a skill in your data model, and how is it versioned?
  2. Which fields are observed, imported, inferred, assessed, and manually entered?
  3. Can each recommendation expose source, date, confidence, and reason?
  4. What evaluation data supports each model in our proposed workflow?
  5. Which metrics are available at our operating threshold and for relevant subgroups?
  6. How are proficiency, confidence, recency, and evidence strength represented separately?
  7. How do employees inspect, correct, and contest data, and where do corrections propagate?
  8. Can our experts override mappings without losing the original evidence and audit trail?
  9. Which taxonomies, models, or thresholds can change without our approval?
  10. What happens when evidence is missing, old, contradictory, or multilingual?
  11. Can we export concepts, relationships, provenance, permissions, and correction history?
  12. Which customers have removed your product, and what did a complete exit require?
  13. Which employment use cases do you classify as high-risk, and what deployer documentation do you provide?
  14. How do you test accessibility, adverse impact, drift, and material model updates?
  15. Which outcome will the pilot improve, and what result would make you recommend that we stop?

The most revealing answers often concern failure and exit. A credible supplier can describe where its product should not be used.

Build, buy, or combine?

Buy the HCM-native layer when one suite already governs identity, jobs, permissions, and the target workflow, and its skills model passes your test. Integration may be worth more than a marginally better specialist algorithm.

Buy a specialist when inference quality, external labour-market data, assessment depth, marketplace activation, or cross-suite neutrality is central—and the vendor can prove that advantage on your data.

Build when the population is bounded, the vocabulary or decision logic embodies distinctive expertise, or vendor abstractions erase essential operational context. Building still requires an owner, evaluation set, governance, integrations, and maintenance; control is not the same as low cost.

Combine when a vendor can provide scale, connectors, workflow, or assessment while your organisation owns the decision contract, core identifiers, evidence policy, labelled evaluation set, acceptance thresholds, and audit history. For many enterprises, this is the durable option: buy replaceable machinery, own the meaning and the accountability.

The 2026 regulatory reality

Regulation follows use, not the label on the licence. Not every skills platform is a high-risk AI system, and the same platform can support low- and high-consequence workflows. In the EU, however, AI used for certain employment purposes—including recruitment, candidate evaluation, promotion, task allocation, and worker monitoring or evaluation—can fall into the AI Act's high-risk category.

As of 27 August 2026, the position is specific: the AI Act is generally applicable, its Article 50 transparency rules have applied since 2 August 2026, and the AI Omnibus moved the Annex III high-risk rules to 2 December 2027. The delay is preparation time, not permission to build an ungoverned system. High-risk obligations include risk management, data quality, logging, documentation, information for deployers, human oversight, accuracy, robustness, and cybersecurity. The Commission also explains that workplace deployers of high-risk systems must inform affected employees and worker representatives before use and assign genuinely equipped human oversight.

GDPR duties continue alongside the AI Act. The European Data Protection Board's guidance remains relevant to profiling and solely automated decisions with legal or similarly significant effects. In New York City, a product within the definition of an automated employment decision tool may trigger Local Law 144, including a recent bias audit, publication, and notice requirements. Scope depends on configuration and use, so require counsel and the vendor to assess each workflow—not “the platform” in the abstract.

Seven signs of vendor overclaiming

  1. Catalogue size is presented as quality. Coverage says little about definition, relevance, duplication, evidence, or local fit.
  2. “Real time” describes a refresh schedule, not observation. Continuously recomputing a weak proxy produces fresher inference, not fresher proof.
  3. An accuracy percentage arrives without a decision threshold. Ask for the task, denominator, benchmark, error costs, and subgroup results.
  4. Inference is displayed as proficiency. Extracting “Python” from text and measuring someone’s ability to use Python are different operations.
  5. “Bias-free”, “objective”, or “ethical AI” is treated as a product property. Demand a defined risk, metric, population, test protocol, result, and remediation process.
  6. A recommendation is reported as an outcome. Matches generated are not roles filled; courses recommended are not skills gained; profiles populated are not workforce agility.
  7. “Single source of truth” means vendor lock-in. A governed source should preserve provenance, disagreement, and exportability. One unchallengeable score is a single point of failure.

Sharp language is not the point of diligence; testability is. Replace every superlative with a measurement request. A serious vendor should welcome the opportunity to prove a bounded claim.

The decision behind the buying decision

The real implementation spans workforce strategy, job architecture, data, AI evaluation, privacy, employee relations, workflow design, and change management. Software can accelerate those parts, but it cannot decide their boundaries or own their consequences.

Buy the platform only after deciding which evidence your organisation will recognise, which decisions it may inform, and who remains accountable when the model is wrong. That is the difference between acquiring another HR database and building skills intelligence.

If you need vendor-neutral support with the decision contract, RFP, evaluation set, or pilot design, get in touch to discuss the selection.

Continue exploring