French Corpus LLM · FINALEADS LLC
Industrial-grade data products for regulated industries.
We design, build and deliver training datasets and bespoke data engineering for finance, public sector, healthcare and other data-intensive verticals — EU AI Act-aligned by design, with full provenance.
Three things we do
Three complementary ways to work with us — pick the one that matches your team’s roadmap.
01 — Data products
Curated training corpora
Off-the-shelf datasets built from authoritative public sources. Cleaned, deduplicated, scored and packaged with the licence chain and audit trail your governance team needs.
2.9M+ documents shipped02 — Custom work
Bespoke data engineering
Pipelines designed for your sources, your schema, your governance rules. From single-source extractors to full curation stacks — you own the IP, we ship the build.
Per-source contract scoped03 — Governance
AI Act readiness
Provenance chains, audit trails and releaseable documentation that satisfy EU AI Act expectations for general-purpose AI training data. Plug into your model card workflow.
AI Act art. 10 readyIndustries we serve
Data-intensive, heavily regulated, document-heavy. The industries where data quality and traceability are not optional.
Finance & Banking Regulation
Regulatory text corpora, prudential and conduct data, market-rules training sets.
ACPR · AMF · BdF · BOFiPPublic Sector & Government
Open-data ingestion, citizen-services language models, legislative analytics.
JORF · LEGI · DILA · data.gouv.frHealthcare & Life Sciences
Clinical-grade corpora, regulatory submissions, scientific literature curation.
clinical-grade trial corporaEnergy & Utilities
Grid & market data, regulatory filings, ESG and sustainability reporting.
CRE · RTE · EUR-Lex energyLegal & Compliance
Case law, contracts and legal-doctrine training data for LegalTech models.
CASS · JADE · CONSTIT · CNILInsurance & Risk
Policy text, claims data structures, risk-modelling reference corpora.
KALI · Solvency II FRWhy customers pick us
Authoritative sources only
Public registers, regulators, central banks and government repositories. No scraped junk, no machine-translated dilution.
JORF · LEGI · CASS · JADE · CONSTIT · KALI · CIRC · CNIL · ACPR · AMF · BdF · BOFiP · DGTrésor · EUR-LexReproducible by design
Every release is versioned, deterministic and re-buildable. Your audit team can verify the artefact end-to-end.
Pipeline SHA-256 : 2fc1de058fd85f3f3eedc73c1f8a89b571b6e… · 2.9M+ JSON-LD audit recordsEU-hosted, AI Act-aligned
Compute and storage hosted in the European Union. Pipeline design built around the AI Act’s data-governance and transparency expectations.
PAdES PKCS#7 self-signed · RSA-4096 · 54-page signed DSD bundleBrowse the blog by topic
Five clusters covering training data, regulatory compliance, and the French vertical.
AI Act & Governance
GDPR pseudonymization, Article 10 documentation, datasheets, audit trails. The compliance moat for LLM training data.
LLM Training Data
Best public datasets in 2026, RAG corpora, SFT vs DPO vs RLHF, instruction tuning, training-data size. What actually trains a model.
Dataset Engineering
Parquet vs JSONL, MinHash LSH deduplication, format trade-offs, dataset versioning. The plumbing that scales.
French NLP & Finance
DILA archives, EUR-Lex FR, ACPR / AMF / CNIL doctrine. The vertical resources you can’t get from generalist English corpora.
Object detection
Training datasets for vision models. Adjacent to LLM data, useful for multi-modal teams.
Compliance & documentation
EU AI Act Article 10, documented
Every release ships with a 38-page Dataset Specification Document, a Source Licences companion, a Glossary, and a cryptographically signed PDF. v1.4.0 — 2,998,388 docs across 27 sources, ~2.66B tokens. All freely downloadable.
📄
Master document
Dataset Specification
The full 10-section AI Act Article 10 / Annex IV documentation: identifier, intended purpose, composition, collection methodology, preparation chain, quality metrics, bias analysis, gap analysis, retention & retraction.
⚖️
Companion
Source Licences
Full enumeration of every licence applicable to source materials: Licence Ouverte 2.0 (Etalab), EU Decision 2011/833/EU, CC-BY-4.0 and Apache 2.0 — with our interpretation of each commercial-reuse permission.
📖
Reference
Glossary
Definitions of every technical, legal, and regulatory term used in the specification — Article 10, Annex III, CELEX, JORF, BOFiP, DILA, ACPR, AMF, MinHash LSH, PROV-O, JSON-LD, and more.
PAdES PKCS#7 self-signed, RSA-4096 · SHA-256 fe19174ba619486c69d1c48f0dc7019b940870bc527788bf6539e48dfde7d272
Latest release fpwc-v1.4.0-2026-05-16 · Specification v1.4 · ACPR and CNIL documents pseudonymized for GDPR safety · 27 sources · ~2.66 billion tokens
From the blog
Latest insights
Field notes on training data, LLM fine-tuning, and AI Act compliance.
BOFiP retrieval — building a French tax-doctrine Q&A system
How to build a question-answering system over BOFiP, the French tax administration’s binding doctrine: why BOFiP is a uniquely structured and frequently-amended…
Comparing French legal LLMs — CamemBERT vs CroissantLLM vs fine-tuning on regulatory text
A decision guide for choosing a base model for French legal and regulatory NLP: where CamemBERT-family encoders win, where CroissantLLM and bilingual…
Building a French RegTech compliance assistant — RAG over DORA, MiCA and the AI Act
A practical build guide for a French RegTech compliance assistant: which EU regulations to ground on (DORA, MiCA, AI Act, plus ESMA…
Tell us what you need
Whether you need an existing dataset, a bespoke pipeline, or compliance advisory — the conversation starts the same way.