French Corpus LLM  ·  FINALEADS LLC

Industrial-grade data products for regulated industries.

We design, build and deliver training datasets and bespoke data engineering for finance, public sector, healthcare and other data-intensive verticals — EU AI Act-aligned by design, with full provenance.

2.9M
Documents
2.6B
Tokens
657k
Jurisprudence rulings
17 sources
Public registers
Last release fpwc-v1.3.0 · 2026-05-16 · live on Hugging Face ↗

Three things we do

Three complementary ways to work with us — pick the one that matches your team’s roadmap.

01 — Data products

Curated training corpora

Off-the-shelf datasets built from authoritative public sources. Cleaned, deduplicated, scored and packaged with the licence chain and audit trail your governance team needs.

2.9M+ documents shipped

02 — Custom work

Bespoke data engineering

Pipelines designed for your sources, your schema, your governance rules. From single-source extractors to full curation stacks — you own the IP, we ship the build.

Learn more →

Per-source contract scoped

03 — Governance

AI Act readiness

Provenance chains, audit trails and releaseable documentation that satisfy EU AI Act expectations for general-purpose AI training data. Plug into your model card workflow.

AI Act art. 10 ready

Industries we serve

Data-intensive, heavily regulated, document-heavy. The industries where data quality and traceability are not optional.

Finance & Banking Regulation

Regulatory text corpora, prudential and conduct data, market-rules training sets.

ACPR · AMF · BdF · BOFiP

Public Sector & Government

Open-data ingestion, citizen-services language models, legislative analytics.

JORF · LEGI · DILA · data.gouv.fr

Healthcare & Life Sciences

Clinical-grade corpora, regulatory submissions, scientific literature curation.

clinical-grade trial corpora

Energy & Utilities

Grid & market data, regulatory filings, ESG and sustainability reporting.

CRE · RTE · EUR-Lex energy

Legal & Compliance

Case law, contracts and legal-doctrine training data for LegalTech models.

CASS · JADE · CONSTIT · CNIL

Insurance & Risk

Policy text, claims data structures, risk-modelling reference corpora.

KALI · Solvency II FR

Why customers pick us

17

Authoritative sources only

Public registers, regulators, central banks and government repositories. No scraped junk, no machine-translated dilution.

JORF · LEGI · CASS · JADE · CONSTIT · KALI · CIRC · CNIL · ACPR · AMF · BdF · BOFiP · DGTrésor · EUR-Lex
2.9M+

Reproducible by design

Every release is versioned, deterministic and re-buildable. Your audit team can verify the artefact end-to-end.

Pipeline SHA-256 : 2fc1de058fd85f3f3eedc73c1f8a89b571b6e… · 2.9M+ JSON-LD audit records
54

EU-hosted, AI Act-aligned

Compute and storage hosted in the European Union. Pipeline design built around the AI Act’s data-governance and transparency expectations.

PAdES PKCS#7 self-signed · RSA-4096 · 54-page signed DSD bundle

Browse the blog by topic

Five clusters covering training data, regulatory compliance, and the French vertical.

AI Act & Governance

GDPR pseudonymization, Article 10 documentation, datasheets, audit trails. The compliance moat for LLM training data.

LLM Training Data

Best public datasets in 2026, RAG corpora, SFT vs DPO vs RLHF, instruction tuning, training-data size. What actually trains a model.

Dataset Engineering

Parquet vs JSONL, MinHash LSH deduplication, format trade-offs, dataset versioning. The plumbing that scales.

French NLP & Finance

DILA archives, EUR-Lex FR, ACPR / AMF / CNIL doctrine. The vertical resources you can’t get from generalist English corpora.

Object detection

Training datasets for vision models. Adjacent to LLM data, useful for multi-modal teams.

Compliance & documentation

EU AI Act Article 10, documented

Every release ships with a 38-page Dataset Specification Document, a Source Licences companion, a Glossary, and a cryptographically signed PDF. v1.4.0 — 2,998,388 docs across 27 sources, ~2.66B tokens. All freely downloadable.

📄

Master document

Dataset Specification

The full 10-section AI Act Article 10 / Annex IV documentation: identifier, intended purpose, composition, collection methodology, preparation chain, quality metrics, bias analysis, gap analysis, retention & retraction.

Read online →

⚖️

Companion

Source Licences

Full enumeration of every licence applicable to source materials: Licence Ouverte 2.0 (Etalab), EU Decision 2011/833/EU, CC-BY-4.0 and Apache 2.0 — with our interpretation of each commercial-reuse permission.

Read online →

📖

Reference

Glossary

Definitions of every technical, legal, and regulatory term used in the specification — Article 10, Annex III, CELEX, JORF, BOFiP, DILA, ACPR, AMF, MinHash LSH, PROV-O, JSON-LD, and more.

Read online →

PAdES PKCS#7 self-signed, RSA-4096 · SHA-256 fe19174ba619486c69d1c48f0dc7019b940870bc527788bf6539e48dfde7d272
Latest release fpwc-v1.4.0-2026-05-16 · Specification v1.4 · ACPR and CNIL documents pseudonymized for GDPR safety · 27 sources · ~2.66 billion tokens

From the blog

Latest insights

Field notes on training data, LLM fine-tuning, and AI Act compliance.

Browse all articles →

Tell us what you need

Whether you need an existing dataset, a bespoke pipeline, or compliance advisory — the conversation starts the same way.