Quantization for LLM Models: Techniques and Trade-Offs

# engineering# oxlo# ai
Quantization for LLM Models: Techniques and Trade-Offsshashank ms

If you are evaluating whether to serve a quantized 4-bit model locally or rely on API access, you need to know exactly where accuracy drops on your ow

If you are evaluating whether to serve a quantized 4-bit model locally or rely on API access, you need to know exactly where accuracy drops on your own data. We will build a small Python evaluator that sends the same long contract clause to a compact model and a flagship model on Oxlo.ai, then diffs the structured outputs to surface the precision cliff.

What you'll need

Step 1: Set up the Oxlo.ai client

I start by importing the OpenAI SDK and pointing it at Oxlo.ai. I will test two endpoints: deepseek-v3.2 as my lightweight stand-in, and llama-3.3-70b as my full-precision reference. Because Oxlo.ai uses flat per-request pricing, I can feed both models a long clause without watching token meters spin. See https://oxlo.ai/pricing for current plan details.

from openai import OpenAI
import json
import sys

client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")

COMPACT_MODEL = "deepseek-v3.2"
FLAGSHIP_MODEL = "llama-3.3-70b"

Step 2: Write the extraction prompt

The system prompt forces strict JSON extraction. I treat any deviation, null confusion, or formatting error as a signal that the model lost precision on the fine print, which is exactly what aggressive quantization risks in production.

SYSTEM_PROMPT = """You are a legal entity extractor. Read the contract clause and return ONLY a JSON object with these keys:
- parties: list of strings
- payment_amount: string or null
- payment_due_date: string or null
- governing_law: string or null
- termination_notice_days: integer or null

If a field is not present, use null. Do not add commentary outside the JSON."""

Step 3: Query the compact model

This helper calls the compact model. In a real deployment, this is analogous to an INT4 quantized 7B model running on a cheap GPU. It is fast, but we suspect it may miss sparse details buried in the text.

def extract_compact(clause: str):
    response = client.chat.completions.create(
        model=COMPACT_MODEL,
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": clause},
        ],
        temperature=0.1,
        max_tokens=512,
    )
    return response.choices[0].message.content

Step 4: Query the flagship model

Now I call the flagship. This acts as my FP16 ground truth. I run the same prompt so the comparison is fair.

def extract_flagship(clause: str):
    response = client.chat.completions.create(
        model=FLAGSHIP_MODEL,
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": clause},
        ],
        temperature=0.1,
        max_tokens=512,
    )
    return response.choices[0].message.content

Step 5: Diff the outputs

I need a robust parser because even small formatting drift can break a downstream pipeline. The diff engine strips markdown fences, parses JSON, and flags every field mismatch. Those mismatches are the quantization risk areas I care about.

def clean_json(text: str):
    t = text.strip()
    if t.startswith("

```"):
        t = t.split("```

")[1].replace("json", "").strip()
    return json.loads(t)

def diff_outputs(compact_text: str, flagship_text: str):
    c = clean_json(compact_text)
    f = clean_json(flagship_text)
    mismatches = []
    for key in f:
        if c.get(key) != f.get(key):
            mismatches.append(
                f"{key}: compact={c.get(key)} | flagship={f.get(key)}"
            )
    return mismatches

Step 6: Wrap the CLI

I tie everything into a small CLI. It prints both results, then lists any divergence. If the compact model matches the flagship on your specific clause, you can be more confident that a quantized local deployment will survive that workload.

def analyze(clause: str):
    print("Querying compact model (deepseek-v3.2)...")
    compact_out = extract_compact(clause)

    print("Querying flagship model (llama-3.3-70b)...")
    flagship_out = extract_flagship(clause)

    print("\n--- Compact Output ---\n", compact_out)
    print("\n--- Flagship Output ---\n", flagship_out)

    mismatches = diff_outputs(compact_out, flagship_out)
    if not mismatches:
        print("\nNo mismatches. The compact model is sufficient for this workload.")
    else:
        print("\nMismatches detected (precision risk):")
        for m in mismatches:
            print(f"  - {m}")

if __name__ == "__main__":
    sample_clause = """
    Party A shall pay Party B the sum of one hundred twenty thousand dollars
    ($120,000) within forty-five (45) days of invoice. This Agreement shall be
    governed by the laws of the State of New York. Either party may terminate
    this Agreement with sixty (60) days prior written notice.
    """
    analyze(sys.argv[1] if len(sys.argv) > 1 else sample_clause)

Run it

I saved the file as quant_eval.py and ran it with the default clause. The compact model returned correct parties and payment_amount, but set termination_notice_days to null while the flagship returned 60. That single miss is the kind of edge-case error that makes quantization risky for contract review. With Oxlo.ai's per-request pricing, running this validation on a hundred sample clauses costs the same per call regardless of whether the prompt is one paragraph or ten pages, so I can build a statistically useful dataset quickly.

$ python quant_eval.py

Querying compact model (deepseek-v3.2)...
Querying flagship model (llama-3.3-70b)...

--- Compact Output ---
{
  "parties": ["Party A", "Party B"],
  "payment_amount": "$120,000",
  "payment_due_date": null,
  "governing_law": "New York",
  "termination_notice_days": null
}

--- Flagship Output ---
{
  "parties": ["Party A", "Party B"],
  "payment_amount": "$120,000",
  "payment_due_date": null,
  "governing_law": "New York",
  "termination_notice_days": 60
}

Mismatches detected (precision risk):
  - termination_notice_days: compact=None | flagship=60

Next steps

Add a router that automatically falls back to the flagship model when mismatches exceed a threshold, or batch-process your own document corpus to compute a task-specific error rate. If the error rate is low enough, you can justify the hardware savings of a quantized local deployment, and if it is high, Oxlo.ai's flat per-request pricing keeps the full-precision API route economical even at long context lengths.