Skip to content

Why a 14B model beats a frontier model on French business law

Le Lion Affaires scores 71.3% on les-audits-affaires against Mistral Medium's 71.2%, with a fraction of the parameters. Here is what the training data did.

Try Le Lion

The result in numbers

Same score. A fraction of the size.

71.3 %Le Lion Affaires
71.2 %Mistral Medium
14 BParameters
−75 %Running cost vs a 70B model

01 · The benchmark

What les-audits-affaires measures

Les-audits-affaires asks French business-law questions the way a lawyer meets them: a clause, a set of facts, a question. A good answer has to cite the applicable text.

It is a hard test for a general model. The language is precise, the sources are national, and an answer that sounds right but cites nothing does not count.

On that test, our 14-billion-parameter model scores 71.3%. Mistral Medium scores 71.2%. The gap is thin; the difference in size is not.

The pipeline

Data in, a smaller model out.

Public lawStatutes, decisions
Written examplesCases, then checked
14B fine-tuneLearns the profession
The generalist testReplayed head to head
↺ If it fails, back to the examples

Simplified view of the training loop.

Score on les-audits-affaires

Le Lion Affaires · 14 B71.3 %
Mistral Medium71.2 %

Same protocol for both models.

02 · How we trained it

Data, not parameters.

  1. IFrame the taskWhich questions, which sources, what a correct answer looks like.
  2. IIGenerate examplesWorked cases written from public law, then checked.
  3. IIIFine-tuneA 14B model learns the profession from those examples.
  4. IVTest against the generalistEvery version is replayed against the model people already use.

Size and cost

Smaller to hold. Cheaper to run.

Parameters, to scale
14 B · Le Lion Affaires
70 B

Five times fewer parameters than a 70-billion model.

Running cost
10070 B model
25Le Lion · 14 B

75% less to run, on hardware an institution already owns.

Square areas to scale (14B vs 70B). Running cost indexed to the 70B model = 100.

In our words

“The smallest model that clears the bar is the right model.”
Mohamad AlhajarParis

03 · Why it matters

A model you can run where the data is

A model of this size runs on hardware a firm or an institution already owns. The documents never leave.

The score matters less than what it allows: the same quality as a frontier model, under your own controls.

Le LionTry it →APIAPI reference →

Read next

All notes →

Winning OpenAI's privacy hackathon at Station F

Dispatches

Building a benchmark with the people who do the work

Benchmarks

Citation is not a feature, it is the whole product

Evaluation

What should your model do?

contact@legml.ai