shashank msIf you are evaluating whether to serve a quantized 4-bit model locally or rely on API access, you need to know exactly where accuracy drops on your ow
If you are evaluating whether to serve a quantized 4-bit model locally or rely on API access, you need to know exactly where accuracy drops on your own data. We will build a small Python evaluator that sends the same long contract clause to a compact model and a flagship model on Oxlo.ai, then diffs the structured outputs to surface the precision cliff.
pip install openaiI start by importing the OpenAI SDK and pointing it at Oxlo.ai. I will test two endpoints: deepseek-v3.2 as my lightweight stand-in, and llama-3.3-70b as my full-precision reference. Because Oxlo.ai uses flat per-request pricing, I can feed both models a long clause without watching token meters spin. See https://oxlo.ai/pricing for current plan details.
from openai import OpenAI
import json
import sys
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
COMPACT_MODEL = "deepseek-v3.2"
FLAGSHIP_MODEL = "llama-3.3-70b"
The system prompt forces strict JSON extraction. I treat any deviation, null confusion, or formatting error as a signal that the model lost precision on the fine print, which is exactly what aggressive quantization risks in production.
SYSTEM_PROMPT = """You are a legal entity extractor. Read the contract clause and return ONLY a JSON object with these keys:
- parties: list of strings
- payment_amount: string or null
- payment_due_date: string or null
- governing_law: string or null
- termination_notice_days: integer or null
If a field is not present, use null. Do not add commentary outside the JSON."""
This helper calls the compact model. In a real deployment, this is analogous to an INT4 quantized 7B model running on a cheap GPU. It is fast, but we suspect it may miss sparse details buried in the text.
def extract_compact(clause: str):
response = client.chat.completions.create(
model=COMPACT_MODEL,
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": clause},
],
temperature=0.1,
max_tokens=512,
)
return response.choices[0].message.content
Now I call the flagship. This acts as my FP16 ground truth. I run the same prompt so the comparison is fair.
def extract_flagship(clause: str):
response = client.chat.completions.create(
model=FLAGSHIP_MODEL,
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": clause},
],
temperature=0.1,
max_tokens=512,
)
return response.choices[0].message.content
I need a robust parser because even small formatting drift can break a downstream pipeline. The diff engine strips markdown fences, parses JSON, and flags every field mismatch. Those mismatches are the quantization risk areas I care about.
def clean_json(text: str):
t = text.strip()
if t.startswith("
```"):
t = t.split("```
")[1].replace("json", "").strip()
return json.loads(t)
def diff_outputs(compact_text: str, flagship_text: str):
c = clean_json(compact_text)
f = clean_json(flagship_text)
mismatches = []
for key in f:
if c.get(key) != f.get(key):
mismatches.append(
f"{key}: compact={c.get(key)} | flagship={f.get(key)}"
)
return mismatches
I tie everything into a small CLI. It prints both results, then lists any divergence. If the compact model matches the flagship on your specific clause, you can be more confident that a quantized local deployment will survive that workload.
def analyze(clause: str):
print("Querying compact model (deepseek-v3.2)...")
compact_out = extract_compact(clause)
print("Querying flagship model (llama-3.3-70b)...")
flagship_out = extract_flagship(clause)
print("\n--- Compact Output ---\n", compact_out)
print("\n--- Flagship Output ---\n", flagship_out)
mismatches = diff_outputs(compact_out, flagship_out)
if not mismatches:
print("\nNo mismatches. The compact model is sufficient for this workload.")
else:
print("\nMismatches detected (precision risk):")
for m in mismatches:
print(f" - {m}")
if __name__ == "__main__":
sample_clause = """
Party A shall pay Party B the sum of one hundred twenty thousand dollars
($120,000) within forty-five (45) days of invoice. This Agreement shall be
governed by the laws of the State of New York. Either party may terminate
this Agreement with sixty (60) days prior written notice.
"""
analyze(sys.argv[1] if len(sys.argv) > 1 else sample_clause)
I saved the file as quant_eval.py and ran it with the default clause. The compact model returned correct parties and payment_amount, but set termination_notice_days to null while the flagship returned 60. That single miss is the kind of edge-case error that makes quantization risky for contract review. With Oxlo.ai's per-request pricing, running this validation on a hundred sample clauses costs the same per call regardless of whether the prompt is one paragraph or ten pages, so I can build a statistically useful dataset quickly.
$ python quant_eval.py
Querying compact model (deepseek-v3.2)...
Querying flagship model (llama-3.3-70b)...
--- Compact Output ---
{
"parties": ["Party A", "Party B"],
"payment_amount": "$120,000",
"payment_due_date": null,
"governing_law": "New York",
"termination_notice_days": null
}
--- Flagship Output ---
{
"parties": ["Party A", "Party B"],
"payment_amount": "$120,000",
"payment_due_date": null,
"governing_law": "New York",
"termination_notice_days": 60
}
Mismatches detected (precision risk):
- termination_notice_days: compact=None | flagship=60
Add a router that automatically falls back to the flagship model when mismatches exceed a threshold, or batch-process your own document corpus to compute a task-specific error rate. If the error rate is low enough, you can justify the hardware savings of a quantized local deployment, and if it is high, Oxlo.ai's flat per-request pricing keeps the full-precision API route economical even at long context lengths.