Why a 14B model beats a frontier model on French business law
Le Lion Affaires scores 71.3% on les-audits-affaires against Mistral Medium's 71.2%, with a fraction of the parameters. Here is what the training data did.
The result in numbers
Same score. A fraction of the size.
01 · The benchmark
What les-audits-affaires measures
Les-audits-affaires asks French business-law questions the way a lawyer meets them: a clause, a set of facts, a question. A good answer has to cite the applicable text.
It is a hard test for a general model. The language is precise, the sources are national, and an answer that sounds right but cites nothing does not count.
On that test, our 14-billion-parameter model scores 71.3%. Mistral Medium scores 71.2%. The gap is thin; the difference in size is not.
The pipeline
Data in, a smaller model out.
Simplified view of the training loop.
Score on les-audits-affaires
Same protocol for both models.
02 · How we trained it
Data, not parameters.
- IFrame the taskWhich questions, which sources, what a correct answer looks like.
- IIGenerate examplesWorked cases written from public law, then checked.
- IIIFine-tuneA 14B model learns the profession from those examples.
- IVTest against the generalistEvery version is replayed against the model people already use.
Size and cost
Smaller to hold. Cheaper to run.
Five times fewer parameters than a 70-billion model.
75% less to run, on hardware an institution already owns.
Square areas to scale (14B vs 70B). Running cost indexed to the 70B model = 100.
In our words
“The smallest model that clears the bar is the right model.”
03 · Why it matters
A model you can run where the data is
A model of this size runs on hardware a firm or an institution already owns. The documents never leave.
The score matters less than what it allows: the same quality as a frontier model, under your own controls.