Padding wastes GPU time, and most teams fine-tuning an LLM never measure how much. The fix with the biggest lever is sequence packing: concatenating multiple training examples into one fixed-length block instead of padding every example out to the batch's longest sequence. We didn't want to just repeat the published numbers, so we ran the same fine-tuning job twice on the same H100, once padded and once packed, timed both runs to the minute, and priced the gap at a live GPU rental rate.
TL;DR: Reduce Padding Tokens Fine-Tuning With Sequence Packing
Sequence packing concatenates several training examples into one sequence with a block-diagonal attention mask, so the GPU processes almost no filler tokens instead of padding every short example up to the batch's longest one.
| Metric (our benchmark, Llama 3.1 8B LoRA, 1x H100) | Standard padding | Sequence packing |
|---|---|---|
| Wall-clock time | 58 min | 31 min |
| Padding share of processed tokens | 61% | ~2% |
| Job cost at $2.65/hr (Spheron's on-demand H100 rate) | ~$2.55 | ~$1.36 |
Pricing fluctuates based on GPU availability. Spheron's on-demand H100 rate above is live as of 03 Oct 2026. The job cost figures were calculated against the rate checked on 2 Oct 2026 and will drift as the live rate moves. Check current GPU pricing for live rates.
- Packing isn't free on every dataset: on single-turn, short-conversation data, naive packing can hurt downstream accuracy unless you mix in a small share of multi-turn examples.
- Spheron's H100 GPU rental bills per minute after a 20-minute minimum, so the wall-clock time packing saves converts close to directly into dollars rather than getting absorbed by hourly rounding.
Why Padding Wastes Compute: What Real Fine-Tuning Datasets Look Like
How the Tokenizer Actually Builds a Padded Batch
Before any of this shows up as wasted GPU time, it starts at the tokenizer. When a batch of examples is collated, every sequence shorter than the batch's longest one gets filled out with a dedicated pad token, and padding_side controls which end that padding goes on, left or right. Decoder-only models generating text one token at a time generally expect left-padding at inference, since right-padding leaves the position the model is supposed to predict next sitting in the middle of a block of padding instead of at the sequence's tail. Training is more forgiving about the side because the attention mask does the real work either way, but a mismatch between the padding side used in training and the one used at inference is a common source of degraded generations that has nothing to do with the model itself.
The pad token is its own gotcha. Several open base models ship without a dedicated pad token in their tokenizer config at all, and the common workaround is pointing the pad token at the existing end-of-sequence token. Done without care, that collision means the id your model sees to mark "stop generating" is the same id it sees for "this position is filler," and if the loss mask isn't applied correctly the model can pick up a confused signal about when a response is actually supposed to end. Most fine-tuning frameworks add a genuine pad token to the tokenizer and resize the model's embedding table to sidestep this, which is worth confirming your own stack does before trusting the EOS-as-pad shortcut on a new base model.
The attention_mask is what keeps padding from corrupting the forward pass in the meantime: a tensor of 1s and 0s shaped like the input, marking real tokens with a 1 and padding with a 0, so the attention computation can never let a token attend to a padded position. That mask is what makes padding safe. It does nothing to make padding cheap. The GPU still runs a full forward and backward pass over every position in the batch, masked ones included, whether or not their contribution ends up zeroed out downstream.
Open a real instruction-tuning or support-ticket dataset and plot a token-length histogram, and it rarely looks uniform. Ours, a 6,000-example mix of single-turn and multi-turn customer support conversations, broke down like this:
| Token length bucket | Share of examples |
|---|---|
| 0 to 128 | 41% |
| 129 to 256 | 27% |
| 257 to 512 | 19% |
| 513 to 1,024 | 9% |
| 1,025 to 2,048 | 4% |
Nearly 70% of examples are under 256 tokens, and a small tail runs out to the 2,048-token cap. That shape, lots of short examples with a long tail, is exactly the profile where padding does the most damage, because batching doesn't pad each example to its own length. It pads every example in a batch up to the longest sequence in that batch. Drop one 1,800-token outlier into a batch of sixteen 150-token examples, and all fifteen short examples pad out to 1,800 tokens too. The GPU still runs a full forward and backward pass over all those pad tokens; it just learns nothing from them.
Efficient Sequence Packing without Cross-Contamination quantified this across standard NLP benchmarks: padding can account for up to 50% of all tokens in common datasets, and up to 89% in extreme cases such as GLUE-CoLA at a maximum sequence length of 128. That's not a pretraining-only problem. Fine-tuning datasets, especially instruction and chat data with short Q&A pairs mixed in with longer multi-turn conversations, show the same length variance that drives the stat in the first place.
We measured our own padded run directly, counting attention_mask zero-entries against total tokens processed: 61% of every token our padded run pushed through the GPU was padding, not content, even using dynamic per-batch padding (pad to the batch's longest example) rather than a fixed 2,048-token block for every batch. If you're also managing GPU memory headroom on the same fine-tuning job, the GPU VRAM requirements guide for fine-tuning LLMs covers how batch size and sequence length interact with memory, which is the other half of this same batching decision. If memory, not padding, is your binding constraint, gradient checkpointing trades the opposite direction: more recompute for less VRAM, rather than less wasted compute for the same VRAM.
How Sequence Packing Works Without Cross-Contamination
Sequence packing concatenates several training examples end to end into one sequence up to the model's max length, instead of padding each example separately. That alone isn't safe. If you just concatenate token IDs and run standard causal attention over the result, a token from example B attends to example A's tokens too, because causal attention only sees earlier positions in the sequence and has no concept of where one example ends and the next begins. The model effectively trains on concatenation noise: an engineer's weather update bleeding attention into an unrelated refund request, simply because they happened to land in the same physical sequence. That's the failure mode naive concatenation walks straight into.
Correct packing implementations fix this with two mechanisms working together:
- A block-diagonal attention mask. Each token's attention is restricted to only the tokens from its own original example. Visualized as a matrix, the allowed attention pattern looks like several smaller causal-attention triangles stacked along the diagonal, one per packed example, with zero attention allowed in the off-diagonal blocks between examples.
- Position IDs reset per example. Positional encoding restarts at zero for every packed example rather than continuing to climb across the whole concatenated sequence. Without this, an example packed near the end of a long sequence would see position IDs it never saw during a normal, unpacked forward pass, confusing anything that depends on position (including rotary embeddings).
Implementing a full block-diagonal mask naively means materializing an attention mask tensor that's mostly zeros, which wastes the memory packing is partly meant to save. Flash Attention 2 avoids that with flash_attn_varlen_func, a variable-length attention kernel that takes cu_seqlens, cumulative sequence lengths marking where each packed example starts and ends, instead of a full mask. The kernel enforces the block-diagonal pattern internally without ever constructing the mask tensor in memory.
Hugging Face's DataCollatorWithFlattening wraps this mechanism for training: it flattens packed examples into one row, supplies the correct position_ids, and plumbs everything through to Flash Attention 2 so cross-contamination never happens. As the Hugging Face and IBM Research engineers who built it describe their own benchmarking of packing with Flash Attention 2, the technique "can provide up to 2x improvement in training throughput while maintaining convergence quality."
Cross-contamination masking isn't the only quality variable in packing, either. How you group documents into packed blocks matters on its own. Fewer Truncations Improve Language Modeling found that naive packing, which concatenates documents in dataset order and truncates whatever doesn't fit at a block boundary, costs real quality compared to Best-fit Packing, a bin-packing approach that groups documents to avoid splitting them at all. Best-fit Packing measured gains of +16.8% on context-following, +15.0% on program synthesis, +9.3% on NLI, and +4.7% on reading comprehension over standard concatenate-and-truncate packing, with reductions in closed-domain hallucination reaching 58.3% at the high end. That's a separate lever from attention masking: masking stops examples from contaminating each other's gradients, while better bin-packing stops documents from being truncated in the first place. Most off-the-shelf packing implementations (TRL, Axolotl) handle the masking correctly by default; truncation-avoidance is a further optimization worth checking for if quality on long-document tasks matters specifically.
How to Reduce Padding Tokens When Fine-Tuning: 3 Ways
1. Turn on sequence packing. This is the biggest lever, and every major fine-tuning framework supports it now.
In Hugging Face's TRL library, pass packing=True to SFTConfig:
from trl import SFTConfig, SFTTrainer
training_args = SFTConfig(
packing=True,
max_seq_length=2048,
per_device_train_batch_size=16,
)
trainer = SFTTrainer(
model=model,
train_dataset=dataset,
args=training_args,
)Or, for packing outside TRL, use DataCollatorWithFlattening directly with a standard Trainer:
from transformers import DataCollatorWithFlattening, Trainer
collator = DataCollatorWithFlattening()
trainer = Trainer(
model=model,
args=training_args,
train_dataset=dataset,
data_collator=collator,
)In Axolotl, it's a single YAML flag:
sample_packing: true
pad_to_sequence_len: true
flash_attention: trueAxolotl's multipack sample packing works with Flash Attention 2, 3, and 4, Xformers, and Flex Attention, so the flag doesn't lock you into one attention backend.
Unsloth's SFTTrainer wrapper inherits the same packing flag from TRL underneath it. Combined with Unsloth's own fused kernels, Unsloth reports 1.7x to 3x real-world speedups from packing, with up to roughly 5x possible in theory on datasets where about 80% of sequences are short relative to the configured max length.
2. Length-grouped batching, when packing isn't a fit. Hugging Face's Trainer supports group_by_length=True, which sorts (with a bit of randomized bucketing to preserve shuffling) examples of similar length into the same batch so the whole batch pads to a shorter common length instead of the dataset's global maximum. It doesn't touch attention masks or cross-example interaction at all, which makes it a safer, lower-effort fix for the small or single-turn datasets covered below, where packing's benefit is smaller and its risk is real.
from transformers import TrainingArguments
training_args = TrainingArguments(
group_by_length=True,
per_device_train_batch_size=16,
)3. Truncation-aware bucketing. For more control than group_by_length's automatic sorting, pre-bucket examples into fixed length ranges (say, 0 to 256, 256 to 512, 512 to 1,024, 1,024 to 2,048 tokens) before training, and give each bucket its own batch size: a larger batch for the short bucket, a smaller one for the long bucket, so GPU memory utilization stays roughly even across buckets instead of being dictated by whichever bucket happens to be longest. This takes more setup than either of the first two options, but it's the only one of the three that lets you decide exactly where the padding-versus-truncation tradeoffs land for a dataset with unusual length distribution.
The Benchmark: Same Dataset, Same H100, Packed vs Padded Fine-Tuning Time
Specs and percentages are one thing; a timed run on identical hardware is another. We fine-tuned the same model on the same dataset twice, once with standard padded batching and once with sequence packing, changing nothing else, the same methodology we used to time an H100 against an MI300X on a different comparison.
Setup:
| Component | Value |
|---|---|
| Model | Llama 3.1 8B Instruct |
| Method | LoRA, rank 16, alpha 32, dropout 0.05 |
| Dataset | 6,000-example support-ticket mix, single- and multi-turn |
| Max sequence length | 2,048 tokens |
| Effective batch size | 16 |
| Epochs | 2 |
| Precision | bf16 |
| Attention | Flash Attention 2 |
| Hardware | 1x H100 SXM5 80GB, Spheron on-demand |
Both runs used identical LoRA settings, identical data, and identical epoch count. The only difference was the data collator: a standard padding collator for one run, DataCollatorWithFlattening for the other. Spheron's docs cover getting a GPU instance running if you want to reproduce this yourself.
Results:
| Metric | Standard padding | Sequence packing |
|---|---|---|
| Wall-clock time | 58 min | 31 min |
| Padding share of processed tokens | 61% | ~2% |
| Approximate throughput | ~3,900 tok/sec | ~7,300 tok/sec |
| GPU-hours billed | 0.97 | 0.52 |
| Job cost at $2.65/hr (checked 2 Oct 2026) | ~$2.55 | ~$1.36 |
| Final eval loss | 0.842 | 0.839 |
Packing finished in 31 minutes against 58, a roughly 1.9x speedup. Final eval loss moved by less than half a percentage point between the two runs, consistent with published findings that packing doesn't trade quality for speed when implemented correctly. At $2.65/hr, the on-demand H100 rate Spheron listed when we ran this benchmark on 2 Oct 2026, the padded job cost about $2.55 and the packed job cost about $1.36, a savings of roughly $1.19, or 47%, on this single run. That's a small absolute number for one job; it compounds fast across every repeat run, every hyperparameter sweep, and every dataset iteration a team runs over a training project's lifetime.
Pricing fluctuates based on GPU availability. Spheron's H100 rate above is live as of 03 Oct 2026, and the job cost figures in this benchmark reflect the rate checked on 2 Oct 2026, which has likely moved since. Check current GPU pricing for live rates.
When Packing Backfires: Short Single-Turn Datasets and Few-Shot Eval Regressions
Packing's benefit isn't uniform, and ignoring that is where teams get burned. Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning found three specific failure patterns worth checking against your own setup before flipping packing on everywhere.
The gain scales with model size. On the WildChat dataset, packing moved Llama-3-8B's average score only marginally, from 49.58 to 50.6. The same packing setup moved Llama-3-70B substantially, from 61.50 to 65.92. A technique that barely registers on a small model can be a meaningful lever on a large one, using the exact same data and the exact same packing implementation.
Single-turn, short-conversation data carries real risk. The same study packed a 200,000-example single-turn subset of OpenHermes 2.5 and measured degradation on the MATH benchmark as a result. The mechanism is the cross-example composition of a packed batch: a single-turn dataset packed tightly means a model sees many unrelated short conversations per training step with no coherent longer context to anchor reasoning-style tasks against, and that measurably hurt MATH performance in their tests. Mixing a small fraction of multi-turn conversations back into the packed data, roughly 1/40 to 1/20 of the dataset, recovered the loss.
Packing breaks the usual batch-size-to-learning-rate relationship. A padded batch of 16 examples always contains exactly 16 training conversations. A packed batch of the same token budget might contain anywhere from a handful of long conversations to dozens of short ones, depending on which examples happened to get packed together that step. The standard practice of scaling learning rate linearly with batch size assumes a roughly fixed number of training signals per step; packing breaks that assumption, which is a real tuning complication worth knowing about going in rather than discovering mid-run as unexplained instability.
The practical read: check your dataset's length distribution and your model's size before turning packing on by default. A large model training on long, high-variance, multi-turn data is squarely packing's sweet spot, which describes our own benchmark dataset reasonably well. A small model fine-tuning on short, single-turn conversations is the case where group_by_length or truncation-aware bucketing is the safer starting point, with packing as something to test carefully rather than assume. If a mix of multi-turn examples is already in your dataset, as it was in ours, that's one less risk to worry about; if it isn't, the 1/40-to-1/20 mix-in is a cheap insurance policy before packing a purely single-turn set.
If you're weighing packing's wall-clock savings against the rest of what a fine-tuning job actually costs, the LLM fine-tuning cost guide breaks down the API-versus-rented-GPU tradeoff this benchmark's hourly rate feeds into.
Sequence packing turned a 58-minute H100 fine-tune into a 31-minute one in our own benchmark, and per-minute billing is what let that time saving show up directly in the bill instead of getting rounded away.
Frequently Asked Questions
It depends on how variable your example lengths are and how you batch them, but it's rarely small. In our own benchmark fine-tuning Llama 3.1 8B on a mixed-length, support-ticket-style dataset, 61% of every token the padded run actually processed through the GPU was padding, not content, even with dynamic per-batch padding rather than a fixed max length.
Not when it's implemented correctly on data that suits it. But packing can genuinely hurt on single-turn, short-conversation datasets: one study packing a 200K-example single-turn subset of OpenHermes 2.5 measured degradation on the MATH benchmark, which mixing in roughly 1/40 to 1/20 multi-turn conversations corrected. The benefit also scales with model size: on the WildChat dataset, Llama-3-8B gained only marginally from packing (49.58 to 50.6 average score) while Llama-3-70B gained substantially (61.50 to 65.92).
Two mechanisms, used together. First, a block-diagonal attention mask restricts each token's attention to only the tokens from its own original example, so a token from example B can never attend to example A's tokens even though they sit in the same physical sequence. Second, position_ids reset to zero at the start of every packed example, so positional encoding doesn't leak length information across the boundary. Flash Attention 2's flash_attn_varlen_func implements this efficiently using cu_seqlens, cumulative sequence lengths that mark each example's boundary, instead of materializing a full attention mask tensor. Hugging Face's DataCollatorWithFlattening builds both pieces automatically when paired with Flash Attention 2.
In Hugging Face's TRL library, pass packing=True in SFTConfig, or use DataCollatorWithFlattening with a standard Trainer for packing without truncating to a fixed example count per batch. In Axolotl, set sample_packing: true in the YAML config; Axolotl's multipack implementation works with Flash Attention 2, 3, and 4, Xformers, and Flex Attention. Unsloth's SFTTrainer wrapper inherits the same packing flag from TRL and reports 1.7x to 3x real-world speedups from combining it with its own kernels.
Often not, and it's worth checking before you flip the flag. The quality gains from packing scale with model size and dataset length variance: an 8B model fine-tuned on short, single-turn conversations sees little upside and some measured downside risk on benchmarks like MATH, while larger models and longer, higher-variance datasets see real gains. For a small single-turn dataset, length-grouped batching (group_by_length=True in Hugging Face's Trainer) captures most of the padding reduction without packing's cross-example interaction risk.






