← Latest reporting

An AI job-matching tool needs a service contract before it needs a ranking model

An OECD report proposes AI-supported matching for employment services in Belgium and Greece while stressing fragmented data, human judgement and gradual deployment. The first design artefact should define the service decision, evidence and appeal path.

Skills Systems and HR TechPolicy, Standards and Governance
A hand-drawn landscape shows two travellers facing separated data islands, several labelled-looking but unreadable service paths and one unfinished bridge.
Conceptual illustration generated with AI under editorial direction; it does not depict a real event.

What happened

On 24 September the OECD published a 182-page report, funded through the EU Technical Support Instrument, on digital employment support in Belgium and Greece. It recommends linked administrative data, gradual AI deployment, human judgement, staff training, monitoring and evaluation.

Why it matters

A matching score is only one component of a public service. Without an explicit purpose, data boundary, counsellor workflow, explanation, contestability and outcome measure, technical accuracy can improve while access, job quality or accountability deteriorates.

The OECD's 182-page report, published on 24 September, examines how Belgium and Greece could improve employment and social services through linked administrative data and AI-supported tools. The project was funded through the EU Technical Support Instrument and implemented with the European Commission. It is a design and policy study, not an impact evaluation of a deployed matching system.

That distinction matters. The report describes fragmented data, uneven digital maturity and governance responsibilities across institutions. It says AI should support rather than replace human judgement and recommends gradual deployment, user feedback, training, safeguards, monitoring and evaluation. These are service-design requirements. They cannot be bolted onto a ranking model after procurement.

Define the decision the service is allowed to make

Start with a one-page service contract. State who the user is, which decision the tool informs, which decisions remain with a counsellor, which data fields are permitted and what a person can do when the result is wrong. Separate job discovery, eligibility, referral, prioritisation and sanction: they have different stakes and should not share one score or review path. A tool that suggests vacancies can tolerate a different error profile from one that changes access to support.

The contract should name the target outcome. Clicks, completed profiles and recommendation acceptance are process measures. Employment entry, retention, earnings, job quality, access for disadvantaged groups and counsellor workload are outcomes or balancing measures. None alone establishes success. For example, faster placement may be paired with poorer job stability, while more counsellor discretion may improve exceptions but increase inconsistency. Define the minimum set before model selection.

Build evidence around the workflow

Create a baseline using the current service, then pilot the AI component in a bounded geography or claimant group. Randomisation may not always be feasible, but a phased rollout can still support comparison if eligibility, labour-market conditions and concurrent policy changes are recorded. Measure who receives recommendations, who acts on them, who is filtered out and how often counsellors override the system. Sample the reasons for overrides rather than treating them as noise.

Data quality should be assessed by decision purpose, not only completeness. A stale occupation code may be harmless for broad exploration and harmful for an eligibility decision. Record provenance, update frequency, lawful basis, known coverage gaps and the institution accountable for correction. Users need a plain-language explanation and a route to challenge material errors without first proving how the model works.

The strongest counterargument is that a detailed service contract may slow experimentation. A short, versioned contract does the opposite: it lets teams test a narrow capability without implying that the whole service is automated. It also makes stopping conditions explicit. Pause expansion when outcome gaps widen, appeal volume rises, data drift exceeds tolerance or counsellors create workarounds to compensate for unusable recommendations.

Governance should include the people operating and receiving the service. Ask counsellors and jobseekers to review examples before launch, then publish a change log for material revisions to ranking logic, data sources and decision rights. Audit samples should include people who received no recommendation, not only accepted matches. Otherwise the evidence base will systematically miss exclusion. Procurement terms should preserve access to logs, error analysis and independent evaluation after the model or vendor changes.

The immediate decision for an employment service is therefore not which matching model leads a benchmark. It is whether the service can specify purpose, decision rights, data responsibility, appeal and outcome evidence. Use the AI exposure explorer to frame task-level changes for counsellors, then test the tool inside that operating model. Ranking quality becomes decision-useful only after the surrounding service can explain, monitor and reverse its effects.