Delph-AIDelph-AI

How Our AI Works

Last Updated: July 16, 2026

This document is only available in English.

EU AI Act Article 50 — Transparency obligations

System Overview

Delph-AI is an AI-assisted screening tool for systematic reviews. It uses multiple large language models (LLMs) to evaluate bibliographic records against user-defined inclusion/exclusion criteria.

AI Models Used

Delph-AI uses up to 40 AI models from 8 providers (OpenAI, Anthropic, Google, Mistral, xAI, and others). Each model evaluates each record independently. The final decision is based on the agreement rate across all selected models — no single model decides alone.

How Decisions Are Made

For each record and each criterion, every selected AI model is requested to return a TRUE/FALSE judgment. A response may fail or be excluded if it cannot be parsed into a valid judgment. A separate textual explanation is generated when disagreement triggers the Delphi reasoning and review process; when the valid Round 0 responses are unanimous, the Service may finalize the classification without generating a separate explanation. The agreement rate (0–1) quantifies consensus among valid responses. The user sets the threshold for inclusion. All decisions are fully auditable and version-controlled.

Agreement Rate

The Agreement Rate (0–1) measures consensus among AI Models for a given record and criterion. It is calculated using only valid, successfully parsed include or exclude responses. Model responses that cannot be parsed into a valid judgment are excluded from both the numerator and denominator.

An Agreement Rate of 1.0 indicates unanimity among valid parsed evaluations. It does not necessarily mean that every selected AI Model returned a usable response. Review the number of valid and failed model responses alongside the Agreement Rate.

No-Consensus Results

A no_consensus result occurs when the pipeline completes its evaluation and valid model responses remain tied after all available Delphi rounds — neither include nor exclude achieved the required threshold.

Model unavailability (due to API outages, quota limits, or persistent technical errors) does not produce a no_consensus result. Instead, the Screening is placed in a stalled state and remains in processing status. Delph-AI will contact you to offer a resolution, such as adjusting the model selection or retrying when the provider becomes available. If no resolution is accepted or agreed upon, the Screening remains in its stalled state and records already debited are not re-credited — see our Terms of Service for the applicable re-credit policy.

A no_consensus result requires manual review and should not be treated as a completed include or exclude recommendation.

Model Explanations

Textual model explanations are generated when the Delphi process enters a reasoning or revision round due to disagreement among models. If models reach unanimity in the initial evaluation round, the Service may return the classification without a separate textual justification for that record.

Limitations

  • AI models may produce incorrect judgments on individual records
  • The consensus mechanism mitigates but does not eliminate errors
  • The human researcher always has the final decision
  • Results depend on the quality of the criteria defined by the user

Data Usage

Your bibliographic data is processed only for the purpose of screening.

Your bibliographic data is not used by Delph-AI to train any AI model. We use contractual restrictions, provider settings, and routing controls to prevent training where available and verified for each provider. Only titles and abstracts of academic publications are sent to AI models — never your personal information. For the current no-training status per provider, see our Sub-Processors page.

For full details about how we handle your data, see our Privacy Policy. For a list of our AI providers and their data processing practices, see our Sub-Processors page.

Regulatory Transparency

Delph-AI is classified as a limited-risk AI system under the EU Artificial Intelligence Act (Regulation (EU) 2024/1689). We are not classified as high-risk because our system evaluates academic publications, not people, and does not make decisions with legal or significant effects on individuals.

In compliance with Article 50 of the EU AI Act, we provide this transparency page to inform you about how our AI system works, what data it processes, and its limitations. Delph-AI is an AI-assisted screening service. Screening classifications and model explanations are generated by AI Models. Users remain responsible for reviewing results and making final inclusion or exclusion decisions.