Skip to content

Portfolio · Legal operations

Legal invoices, read in fullrather than in sample.

Law firms and outside counsel bill corporate legal departments and insurers in thousands of line items a month, against billing guidelines nobody has time to check line by line. The model reads every one: it classifies the item and applies the service-level and guideline code, so an invoice arrives at approval already reviewed. Programmes of this kind take between 6% and 11% off annual legal spend.

Legal spend is controlled one line at a time, or it is not controlled.
Sector
Legal operations — outside counsel spend and invoice compliance
Model
BERT transformer, fine-tuned on line-item text
Coverage
Every line item, rather than a sampled subset
Outcome
6–11% average annual reduction in legal spend
01

The unit of review

/ 03

The unit is one line, and there are thousands of them.

A legal invoice is not reviewed as a document. It is reviewed as a list: each line a narrative written by a fee earner, each one either inside the client’s billing guidelines or outside them, and the difference frequently a matter of a few words. Manual review handles that by sampling — which finds the pattern eventually, and pays for everything in the meantime. Classifying every line changes what review can be: complete rather than representative.

  1. 01

    Complete, not representative

    Sampling finds the pattern eventually and pays for everything it did not sample in the meantime.

  2. 02

    The narrative is the evidence

    What separates a compliant line from a rejected one is usually a few words in the description.

  3. 03

    A code, not an opinion

    The output is the service-level or guideline code the review process already runs on.

Thousands of lines a month, each one a sentence somebody wrote and somebody else has to judge.
The evidence is the narrative text, which is why this is a language problem and not a rules engine.
A classification is only useful once it is a code the billing system already understands.
02

The model

/ 03

A transformer, because the signal is in the wording.

Line narratives are short, domain-specific and adversarially similar: two lines describing the same hour of work can fall on opposite sides of a guideline. A fine-tuned BERT classifier handles that where keyword rules do not, reading the description in context and assigning both the classification and the code that follows from it. Built in Python on PyTorch over the usual preparation, by a team of three.

  1. 01

    Context beats keywords

    A rule on words matches the wrong lines as reliably as it matches the right ones.

  2. 02

    Trained on the real set

    Fine-tuned against actual reviewed invoices, which is where the guideline is really written down.

  3. 03

    Small team, narrow scope

    Three people, one model, one job — the scope is why it reached production.

Fine-tuned on the client’s own line items, because a general model does not know these guidelines.
One line in, one code out — and the borderline cases are exactly where the value sits.
03

What it returns

/ 03

Between six and eleven per cent, every year.

The point of reviewing every line is not the catch rate on any one invoice; it is what happens to the invoices that arrive afterwards. A billing programme that reviews in full and codes consistently takes 6–11% off annual legal spend, and it does it without the cost to a client relationship that a manual, arbitrary and occasionally wrong challenge tends to carry.

  1. 01

    Behaviour follows review

    Firms bill differently once they know every line is read, which is where most of the saving lives.

  2. 02

    Consistent across firms

    One guideline, applied the same way to everybody, at any volume.

  3. 03

    Auditable by line

    Every decision is attached to the line it was made about, so a challenge is a conversation about text.

6–11% a year — and the larger part of it comes from the invoices sent after the first review, not before.
The same guideline applied the same way to every firm, which is the part a person cannot promise.
Spend becomes a series somebody manages rather than a total somebody explains.

The last word

This shows how Famysys builds a model into a process that already exists: take the unit the reviewers actually work in, classify all of it rather than a sample, and return a code the billing system already understands. If your spend controls depend on somebody having time this month, they are not controls.

Built in

  • BERT transformer
  • Line-item classification
  • Guideline code assignment
  • Invoice compliance
  • PyTorch
  • Full-coverage review
  • Spend analytics

All projects

Let's build what's next

Technology alone doesn't transform businesses.The right partnership does.

Modernizing systems, building a new product, or exploring AI-driven transformation — Famysys helps you move forward with confidence.

  • SOC2 Type II Compliant
  • Strict Commercial NDA
  • Zero Lock-In Guarantee

Famysys

We Engineer Clarity

Scan to Connect

Instant digital business card & WhatsApp link