+1 (726) 227-4060

Fine-Tuning on Private Data: Differential Privacy with Opacus and Federated LoRA in PyTorch

Every enterprise PyTorch engagement eventually hits the same wall: the data that would make the model good is the data legal will not let you train on. Claims notes, support transcripts with account numbers, clinical records, internal payroll tables. The usual answer — "we'll scrub PII first" — is necessary but not sufficient, because a fine-tuned model can memorize and later regurgitate rare training examples even after a scrubbing pass. Regulators and enterprise security reviewers have caught up to this, and "we removed names" is no longer an argument that survives a data protection impact assessment.

This tutorial covers the two PyTorch-native mechanisms that are arguments: differentially private fine-tuning with Opacus, and federated fine-tuning, where the data never leaves its silo at all. We cover when each one is the right call, what they cost you in accuracy and wall-clock time, and the specific configuration that makes DP-LoRA workable on a single node.

First: what problem are you actually solving?

These are different threat models and people conflate them constantly. Before writing any code, sort your engagement into one of three buckets.

ConcernRight toolWrong tool
Model might emit a training record verbatim to a userDifferential privacy (DP-SGD)Federated learning alone
Data cannot legally leave a hospital / region / customer tenantFederated learning, or per-tenant adaptersDP alone
You just don't want raw PII in your artifact storeRedaction + tokenization pipelineEither of the above

Federated learning solves data movement. Differential privacy solves memorization. If your risk register contains both rows, you need both, and the combination is the standard "federated learning with a DP aggregator" setup at the end of this post.

And if your only real concern is the third row, do not reach for DP. Redaction plus a tightly scoped dataset is cheaper, faster and easier to explain to an auditor. We have talked more than one client out of a DP program that would have cost them six accuracy points to solve a problem a regex and a retention policy already solved.

Differentially private fine-tuning with Opacus

DP-SGD changes the optimizer, not the model. Two modifications to each step:

  1. Per-sample gradient clipping. Compute the gradient for each example individually, clip its L2 norm to C, then sum. This bounds how much any single record can move the weights.
  2. Gaussian noise. Add noise with standard deviation sigma * C to the summed gradient before the optimizer step.

Together these give you an (epsilon, delta)-differential privacy guarantee: the resulting weights are provably almost as likely to arise from a dataset without any given record as from one with it. epsilon is the privacy budget — smaller is stronger. In practice, enterprise fine-tuning lands between epsilon = 1 (strict, common in health and finance) and epsilon = 8 (weak but still defensible against extraction attacks), with delta set to roughly 1 / (10 * N) for a dataset of N records.

Opacus is the PyTorch library that implements this. The core wiring:

import torch
from opacus import PrivacyEngine
from opacus.utils.batch_memory_manager import BatchMemoryManager

model, optimizer, train_loader = build_model_and_data()  # your usual setup
model.train()

privacy_engine = PrivacyEngine(accountant="prv")

model, optimizer, train_loader = privacy_engine.make_private_with_epsilon(
    module=model,
    optimizer=optimizer,
    data_loader=train_loader,
    target_epsilon=4.0,
    target_delta=1e-5,
    epochs=3,
    max_grad_norm=1.0,          # the clipping bound C
    poisson_sampling=True,      # required for the accounting to be valid
)

with BatchMemoryManager(
    data_loader=train_loader,
    max_physical_batch_size=4,   # what fits on the GPU
    optimizer=optimizer,
) as memory_safe_loader:
    for epoch in range(3):
        for batch in memory_safe_loader:
            optimizer.zero_grad()
            loss = model(**batch).loss
            loss.backward()
            optimizer.step()

print("spent epsilon:", privacy_engine.get_epsilon(delta=1e-5))

Three details in there matter more than they look.

poisson_sampling=True is not optional. The privacy accounting assumes each example is included in a batch independently with probability B / N, not that you shuffled and chunked. Opacus swaps in a UniformWithReplacementSampler to do this. If you disable it because it breaks your dataloader, your reported epsilon is a number with no theorem behind it.

BatchMemoryManager decouples logical from physical batch size. DP-SGD wants large logical batches — a batch of 1024 with noise sigma is far more accurate than 16 batches of 64 each noised — but per-sample gradients are memory hungry. The memory manager accumulates clipped per-sample gradients across micro-batches and only noises and steps at the logical boundary. Set max_physical_batch_size to whatever fits, and set the logical batch size in your sampler as large as your budget allows. This is the single biggest accuracy lever in DP fine-tuning.

The prv accountant is the one to use. The older RDP accountant is looser and will overstate your spent epsilon, costing you accuracy for nothing.

Make it DP-LoRA, not DP-full-fine-tune

Full-parameter DP-SGD on a 7B model is miserable: per-sample gradients for every weight, and noise injected into billions of coordinates. Parameter-efficient fine-tuning fixes both problems at once, and the interaction is genuinely synergistic rather than just convenient — noise is added only to the small adapter, so the signal-to-noise ratio per trainable parameter is far better at the same epsilon.

from peft import LoraConfig, get_peft_model
from opacus.validators import ModuleValidator

base = load_base_model()
lora = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"],
                  lora_dropout=0.0)   # see note below
model = get_peft_model(base, lora)

# Opacus needs every module to support per-sample grads.
errors = ModuleValidator.validate(model, strict=False)
model = ModuleValidator.fix(model)    # e.g. BatchNorm -> GroupNorm

Set lora_dropout=0.0. Dropout inside the adapter interacts badly with per-sample gradient hooks and, more importantly, you are already regularizing with gradient noise; stacking a second stochastic regularizer just costs you convergence.

ModuleValidator exists because some layers are incompatible with per-sample gradients. BatchNorm is the classic offender — it mixes information across examples in a batch, which breaks the per-sample premise outright. ModuleValidator.fix rewrites it to GroupNorm. Run the validator before you start a long run, not after.

Hyperparameters that actually change the outcome

DP fine-tuning inverts several habits from normal training:

  • Batch size: as large as you can afford. Effective batch sizes of 512 to 4096 (via accumulation) are normal. Bigger batch, better signal-to-noise, cheaper privacy per example.
  • Epochs: few. Every pass over the data spends budget. Two to three epochs at a large batch beats ten at a small one.
  • Learning rate: higher than you expect, often 5x to 10x your non-private LoRA rate, because clipping shrinks the effective gradient magnitude.
  • Clipping norm C: low, around 0.1 to 1.0. Counterintuitively, aggressive clipping usually helps — the noise scales with C, so a small C means less noise, and the direction of the gradient survives clipping even when the magnitude does not.
  • Never tune hyperparameters on the private data without counting it. Every run that touches the training set spends budget. In practice teams tune on a public proxy dataset and do a single production run.

Expect to give up something. On typical enterprise text classification and extraction tasks, DP-LoRA at epsilon = 8 lands within about a point of non-private fine-tuning; at epsilon = 1 the gap is usually a few points and occasionally much worse on rare classes. Budget for a 2x to 4x wall-clock slowdown from per-sample gradient computation, and measure the gap explicitly before you promise anyone a number.

Federated fine-tuning: when the data cannot move

Sometimes the blocker is not memorization, it is jurisdiction. Three hospitals, four bank subsidiaries, or a product with per-tenant data residency commitments. Nobody will ship you the raw rows.

The federated pattern: each silo trains locally on its own data, sends only weight updates to a central aggregator, and the aggregator averages them into a new global model (FedAvg). With LoRA, the thing being shipped over the wire is a few tens of megabytes of adapter, not a 14 GB checkpoint — which is what made federated fine-tuning of large models practical in the first place.

A minimal aggregation step, which is all FedAvg really is:

import torch
from collections import OrderedDict

def fedavg(client_states, client_sizes):
    """Weighted average of client LoRA state dicts."""
    total = sum(client_sizes)
    avg = OrderedDict()
    for key in client_states[0]:
        avg[key] = sum(
            state[key].to(torch.float32) * (n / total)
            for state, n in zip(client_states, client_sizes)
        ).to(client_states[0][key].dtype)
    return avg

global_adapter = fedavg(collected_states, [len(d) for d in client_datasets])

Round structure: broadcast global_adapter to each client, each client runs a small number of local epochs, collect, average, repeat for 10 to 50 rounds. Flower (flwr) is the usual framework for the orchestration, retries and client sampling; it integrates with plain PyTorch training loops and with Opacus if you want per-client DP on top.

Three things bite teams here:

  1. Non-IID data. Silos have genuinely different distributions and plain FedAvg can oscillate or converge to something worse than any single silo's local model. FedProx (a proximal term pulling local weights toward the global ones) or server-side momentum (FedAdam) usually fixes it. Always evaluate the global model on each silo's held-out set separately, not just on a pooled average — a global model that is great on your biggest client and useless on the other three is a failed project even if the mean looks fine.
  2. Weight updates leak. Gradient-inversion attacks can reconstruct training samples from raw updates. If the aggregator is not fully trusted, add secure aggregation (updates are masked and only their sum is decryptable) and/or client-level DP noise.
  3. Operations. Version skew, stragglers, and clients that drop mid-round are the real cost of federated learning. Assume the orchestration work exceeds the ML work.

Combining both

The defensible enterprise setup for regulated data is federated rounds where each client applies DP-SGD locally (client-level DP), plus secure aggregation at the server. You get: data never moves, the aggregator cannot read individual updates, and the released model carries a formal bound on what it memorized about any one record.

A pragmatic checklist before you start

  • Write down the threat model in one sentence. If you cannot, you are not ready to pick a mechanism.
  • Confirm a redaction pipeline exists anyway. DP is not a substitute for not logging credit card numbers.
  • Establish the non-private baseline first. You cannot price the privacy cost without it.
  • Pick epsilon with legal and security in the room, before training, and record it in the model card alongside delta, the clipping norm, the accountant and the number of epochs. An epsilon without its accounting assumptions is decoration.
  • Run ModuleValidator and a 50-step smoke run before booking a week of GPU time.
  • Evaluate on rare classes and minority segments specifically. DP noise hurts the tail first, and the tail is often the compliance-relevant part.

Where this goes wrong in production

The most common failure we get called in to fix is not a bad epsilon — it is a pipeline where DP was applied to the fine-tuning step while the retrieval index alongside it contains the same raw documents in plaintext, queryable by any user of the assistant. Differential privacy protects the weights. If your RAG corpus is unfiltered, the model does not need to memorize anything to leak it. Privacy is a property of the whole system, and the retrieval path is usually the weakest link.

The second most common: a privacy budget spent on hyperparameter search, then a final run reported at the per-run epsilon. Budgets compose. Every run counts.


Building on private or regulated data? IntelliSensei's PyTorch consultants design and implement differentially private and federated training pipelines, from threat modelling and epsilon selection through DP-LoRA implementation, secure aggregation and the accuracy evaluation your reviewers will ask for. Get in touch to scope an engagement.