contractguard-clause-classifier

A fine-tuned Legal-BERT sequence classifier that labels a contract clause with one of 40 clause types. It is the first stage of ContractGuard AI, an open-source clause-level contract analysis platform β€” the clause type this model predicts conditions every downstream stage (risk scoring, retrieval, negotiation suggestions).

Not legal advice. This model produces a classification, not a legal conclusion. Its output is intended as decision support for a qualified reviewer, not a substitute for one.

Model summary

Base model nlpaueb/legal-bert-base-uncased (110M parameters)
Task Multi-class sequence classification, 40 clause types
Training data CUAD + LEDGAR (union), 88,115 examples after validation
Version 2.0 (active) β€” supersedes 1.0 (deprecated)
Max sequence length 256 tokens
License Other β€” see License

This is version 2.0 of the classifier. It replaces v1.0, which was trained on LEDGAR alone and never activated in production β€” see Why v2.0, not v1.0 below.

Intended use

Given the text of a single contract clause, predict which of 40 clause types it is (e.g. TERMINATION, INDEMNIFICATION, NON_COMPETE, LIMITATION_OF_LIABILITY). The predicted label and its confidence score feed a downstream risk model, which conditions its own prediction on the clause type, contract type, party role and jurisdiction.

In scope: English-language commercial contracts, segmented into individual clauses before inference.

Out of scope: whole-document classification, non-English contracts, and any of the six clause types listed in Coverage below, which the model was not trained to emit. It also does not detect a missing clause β€” absence is handled by a separate rule engine, since there is no text for a per-clause classifier to act on.

How to use

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("AnuragBodkhe/contractguard-clause-classifier")
model = AutoModelForSequenceClassification.from_pretrained("AnuragBodkhe/contractguard-clause-classifier")

clause = "Either party may terminate this Agreement upon thirty (30) days' written notice."

inputs = tokenizer(clause, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
    logits = model(**inputs).logits

probs = torch.softmax(logits, dim=-1)
label_id = probs.argmax().item()
print(model.config.id2label[label_id], probs[0, label_id].item())

The label set and display names are also published in labels.py in the main repository, which is the single source of truth the rest of the platform mirrors β€” see Taxonomy below.

Training data

Dataset Role
CUAD commercial contract clauses, 21 of 40 label types
LEDGAR SEC-filed contract clauses, 28 of 40 label types

91,144 raw examples were combined and reduced to 88,115 after validation, then split by document (not by clause) into train / validation / test to prevent leakage between splits:

Split Examples
Train 61,483
Validation 13,151
Test 13,481

Taxonomy & coverage

CUAD and LEDGAR each cover a different, overlapping subset of the 40-label taxonomy. Their union covers 34 of 40 (85%):

LEDGAR covers          28 / 40
CUAD covers            21 / 40
UNION                  34 / 40

CUAD is what makes EXCLUSIVITY, LIABILITY, LICENSE_GRANT, NON_COMPETE, NON_SOLICITATION and RENEWAL emittable at all β€” and LIABILITY and NON_COMPETE are exactly the two types whose absence sank v1.0 (see below).

Six labels have no training support in this release and the model cannot produce them: DATA_PROTECTION, FORCE_MAJEURE, MORAL_RIGHTS, PROBATION, SERVICE_COMMITMENT_BOND, SUBCONTRACTING. A label absent from training data is a label the model cannot produce β€” closing this gap requires adding labelled examples and retraining, not a configuration change.

Training procedure

Hyperparameter Value
Epochs 3
Batch size 16 Γ— 2 gradient accumulation
Learning rate 3e-5
Max sequence length 256
Class weighting sqrt-balanced
Precision fp16
Hardware 1Γ— RTX 3050 6 GB (laptop GPU), CUDA 12.6
Wall clock 59 minutes

Evaluation results

Measured on the held-out test split (13,481 examples, split by document):

Metric v1.0 (LEDGAR only) v2.0 (CUAD + LEDGAR)
Accuracy 0.8624 0.9168
Macro F1 0.8592 0.8923
Weighted F1 0.8618 0.9167
Expected calibration error 0.1136 0.0391

The calibration improvement is as consequential as the F1 gain: the downstream UI shows this model's confidence next to every prediction, and v1.0 was materially overconfident.

Why v2.0, not v1.0

v1.0 scored a respectable 0.8592 macro F1 and was still never activated in production, because LEDGAR's 28-label coverage disproportionately excluded the clause types that carry risk. Adding CUAD lifted LIABILITY to F1 0.946 and NON_COMPETE to F1 0.660 β€” both previously unlearnable.

Real-contract validation

Aggregate test-set metrics can hide exactly the failure that matters, so activation was decided on a real document: a UK freelance agreement, analysed once with the deterministic baseline and once with v2.0.

Baseline v2.0
Overall risk score 88 (CRITICAL) 88 (CRITICAL)
Clause types β€” 10 of 11 identical to baseline
Risk findings 11 11, none lost

The one clause-type disagreement (RENEWAL vs. TERM_AND_DURATION) produced the same downstream risk finding either way. Run against v1.0, the same test lost two CRITICAL-severity findings β€” the failure mode aggregate metrics did not surface.

v2.0 also declares the population it was trained on, and flags a real mismatch on this document:

fitted on United States (federal), United States – other state. This contract was submitted under United Kingdom.

fitted on COMMERCIAL, LICENSING, SERVICE, VENDOR contracts; this is a FREELANCE contract.

Both are true, and both are useful things for a reviewer to know before trusting the output.

Limitations and bias

  • Training population. CUAD and LEDGAR skew toward US commercial and SEC-filed agreements. Predictions on contracts from other jurisdictions or contract families (e.g. employment, consumer) should be treated with more caution β€” the model has no way to signal this itself, which is why the serving platform declares the training population alongside every prediction rather than relying on the model to know what it doesn't know.
  • Label coverage. Six of 40 taxonomy labels are un-emittable in this version (see Coverage).
  • Input granularity. The model expects a single segmented clause, not a whole document or an unsegmented paragraph. It was not evaluated on document-level input.
  • Not a legal conclusion. A predicted clause type is a classification under uncertainty, not a determination of legal effect. It is intended to feed a review workflow with a human in the loop.

License

The base model, nlpaueb/legal-bert-base-uncased, is licensed CC-BY-SA-4.0. Training data licenses: CUAD (CC BY 4.0), LEDGAR (subject to the LexGLUE / EDGAR terms). Review those upstream licenses before commercial use in combination with this model's own license terms.

Citation

If you use this model, please cite the base model and the datasets it was fine-tuned on:

@article{chalkidis2020legalbert,
  title={LEGAL-BERT: The Muppets straight out of Law School},
  author={Chalkidis, Ilias and others},
  year={2020}
}
@article{hendrycks2021cuad,
  title={CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review},
  author={Hendrycks, Dan and others},
  year={2021}
}

Part of ContractGuard AI

This model is one of five in the ContractGuard AI pipeline (clause classification, risk classification, legal NER, retrieval embeddings, contract NLI). See the main repository for the full architecture, training results for all five models, and the taxonomy definitions this model uses.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for AnuragBodkhe/contractguard-clause-classifier

Finetuned
(110)
this model

Dataset used to train AnuragBodkhe/contractguard-clause-classifier