Free tools Windows power users keep installed
One-click scans. No signup required.
This tutorial trains a task adapter—not the entire RoBERTa encoder—for binary text classification. It uses the current adapters package, freezes FacebookAI/roberta-base, trains a bottleneck adapter and classification head, evaluates the result, and saves an adapter that can be reloaded with the compatible base model.
What you are building
An adapter is a small trainable module inserted into a pretrained transformer. RoBERTa supplies general language representations; the adapter learns task-specific behavior while the standard encoder weights remain frozen.
Input text
↓
RoBERTa tokenizer
↓
Frozen RoBERTa encoder
↓
Trainable task adapter
↓
Trainable classification head
↓
Class logits
Adapters are modular: several task adapters can use one downloaded base model, and each adapter can be distributed separately. An adapter is not a complete standalone model. Reloading normally also requires the compatible base checkpoint, tokenizer, adapter configuration, and—unless deliberately shared—the prediction head.
The original adapter study reported GLUE results within 0.4 percentage points of full fine-tuning while adding 3.6% task-specific parameters per task in its experimental setup. That is a historical result, not a performance guarantee for your data or current library configuration (original adapter paper).
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Choose the right adaptation method
| Method | Use it when | What is saved |
|---|---|---|
| Classic bottleneck adapter | You want modular task or language adapters, composition, or AdapterHub interoperability. | Adapter configuration and weights, plus a head when needed. |
| LoRA or another PEFT method | Your project already uses PEFT and you prefer low-rank updates, IA3, AdaLoRA, or prefix tuning. | PEFT adapter configuration and weights. |
| Full fine-tuning | Maximum task-specific flexibility matters more than small artifacts and shared-base modularity. | A complete updated model, usually much larger. |
This article uses bottleneck adapters. The current adapters package replaced the older adapter-transformers package while retaining compatibility with previously trained adapter weights (Hugging Face adapter documentation). Transformers also integrates with PEFT (the current documentation lists peft >= 0.19.1), but PEFT and Adapters are different APIs (Transformers PEFT documentation).
Prerequisites and installation
- Python 3.9 or newer.
- PyTorch 2.0 or newer, as listed by the AdapterHub project page (project requirements).
- A labeled dataset with stable training and evaluation splits.
- A GPU for practical datasets; a CPU is sufficient for a small demonstration.
- Disk space for the base model, tokenizer, dataset cache, checkpoints, and adapter output.
python -m venv .venv
source .venv/bin/activate # macOS/Linux
# .venvScriptsactivate # Windows
python -m pip install -U pip
pip install -U adapters datasets evaluate accelerate scikit-learn
Pin and record the versions used for a reproducible project. Package APIs change; run the example against the versions you intend to deploy.
Prepare a labeled dataset
The example uses IMDb sentiment data. Its preprocessing assumes columns named text and label; adapt those names for your data.
from datasets import load_dataset
dataset = load_dataset("imdb")
A CSV dataset can be loaded instead:
dataset = load_dataset(
"csv",
data_files={
"train": "train.csv",
"validation": "validation.csv",
"test": "test.csv",
},
)
- Use integer class IDs beginning at zero unless your setup explicitly handles another label format.
- Keep a validation split for model and hyperparameter decisions; reserve the test split for final reporting.
- Convert string labels to a documented mapping such as
negative → 0andpositive → 1. - For imbalanced classes, report per-class metrics and macro or weighted F1, not accuracy alone.
Load RoBERTa and tokenize the text
from transformers import AutoTokenizer
from adapters import AutoAdapterModel
model_name = "FacebookAI/roberta-base"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoAdapterModel.from_pretrained(model_name)
Use the same base-model identifier for the tokenizer and model. A RoBERTa-base adapter is not automatically compatible with RoBERTa-large, BERT, DeBERTa, or XLM-RoBERTa (RoBERTa documentation).
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsdef preprocess_function(examples):
return tokenizer(
examples["text"],
truncation=True,
max_length=256,
)
tokenized_dataset = dataset.map(
preprocess_function,
batched=True,
remove_columns=["text"],
)
max_length=256 is an example. Longer sequences preserve more context but use more memory and time; shorter sequences may discard useful text. Dynamic batch padding avoids padding every record to the global maximum.
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
For paired inputs, pass both fields:
def preprocess_function(examples):
return tokenizer(
examples["sentence1"],
examples["sentence2"],
truncation=True,
max_length=256,
)
The standard Transformers sequence-classification workflow uses the same truncation, dataset mapping, and dynamic-padding pattern (sequence-classification guide).
Add the adapter and classification head
adapter_name = "sentiment"
head_name = "sentiment"
model.add_adapter(
adapter_name,
config="pfeiffer",
)
model.add_classification_head(
head_name,
num_labels=2,
id2label={
0: "NEGATIVE",
1: "POSITIVE",
},
)
model.active_head = head_name
The pfeiffer configuration selects a common bottleneck architecture. Head signatures have varied across older and newer releases, so confirm the installed version’s add_classification_head() signature. If your release associates a head with the adapter name, use that documented form and activate the resulting head explicitly.
Freeze RoBERTa and train only the adapter
model.train_adapter(adapter_name)
model.set_active_adapters(adapter_name)
train_adapter() freezes ordinary RoBERTa parameters and enables the selected adapter for training. The classification head must also be active and trainable. This is different from merely adding an adapter: forgetting this step can leave the wrong parameters trainable (AdapterHub training documentation).
Audit the result before starting:
def trainable_parameters(model):
total = 0
trainable = 0
for parameter in model.parameters():
count = parameter.numel()
total += count
if parameter.requires_grad:
trainable += count
return trainable, total
trainable, total = trainable_parameters(model)
print(f"Trainable: {trainable:,}")
print(f"Total: {total:,}")
print(f"Percent: {100 * trainable / total:.2f}%")
The percentage depends on adapter architecture and bottleneck size, model size, whether the head or embeddings are trained, and package configuration. The frozen base still occupies memory and participates in forward and backward computation.
Train with AdapterTrainer
import evaluate
import numpy as np
from adapters import AdapterTrainer
from transformers import DataCollatorWithPadding, TrainingArguments
accuracy = evaluate.load("accuracy")
f1 = evaluate.load("f1")
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1)
return {
"accuracy": accuracy.compute(
predictions=predictions, references=labels
)["accuracy"],
"f1": f1.compute(
predictions=predictions,
references=labels,
average="binary",
)["f1"],
}
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)
training_args = TrainingArguments(
output_dir="roberta-sentiment-adapter",
learning_rate=1e-4,
per_device_train_batch_size=16,
per_device_eval_batch_size=16,
num_train_epochs=3,
weight_decay=0.01,
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
report_to="none",
)
trainer = AdapterTrainer(
model=model,
args=training_args,
train_dataset=tokenized_dataset["train"],
eval_dataset=tokenized_dataset["validation"],
processing_class=tokenizer,
data_collator=data_collator,
compute_metrics=compute_metrics,
)
trainer.train()
The learning rate, batch size, epoch count, and sequence length are starting points, not guarantees. Three epochs may underfit or overfit. GPU memory determines feasible batch size. Recent Transformers releases use eval_strategy and processing_class; older releases may require evaluation_strategy and tokenizer. Match the argument names to your pinned versions.
Rank #3
For multiclass classification, use an appropriate macro or weighted F1 average. Multilabel classification needs a different loss, sigmoid outputs, and thresholding; do not reuse this binary setup unchanged.
Evaluate on held-out data
metrics = trainer.evaluate(
eval_dataset=tokenized_dataset["test"]
if "test" in tokenized_dataset
else tokenized_dataset["validation"]
)
print(metrics)
Do not tune repeatedly on the final test set. Inspect a confusion matrix and per-class errors, especially when classes are imbalanced. For serious comparisons, repeat training with controlled seeds and report variation rather than treating one run as definitive.
Save the adapter and tokenizer
model.save_adapter(
"sentiment_adapter",
adapter_name,
with_head=True,
)
tokenizer.save_pretrained("sentiment_adapter")
with_head=True stores the task head with the adapter package. Omitting it can leave classification inference incomplete unless you intentionally maintain a separately shared head. A trainer checkpoint is different: it may also contain optimizer state, scheduler state, trainer state, and intermediate checkpoints. Use a checkpoint to resume training; use the adapter export to share or deploy the trained module.
Record the base model identifier, adapter configuration, library versions, label IDs and names, tokenizer settings, maximum sequence length, dataset provenance, evaluation results, hyperparameters, license, and intended-use limitations. The Hub supports adapter publishing with generated metadata through push_adapter_to_hub() (Hub adapter workflow).
Reload the adapter for inference
import torch
from adapters import AutoAdapterModel
from transformers import AutoTokenizer
base_model = "FacebookAI/roberta-base"
tokenizer = AutoTokenizer.from_pretrained("sentiment_adapter")
inference_model = AutoAdapterModel.from_pretrained(base_model)
inference_model.load_adapter(
"sentiment_adapter",
set_active=True,
)
inference_model.eval()
text = "The product was easy to use and worked well."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
)
with torch.no_grad():
outputs = inference_model(**inputs)
prediction = outputs.logits.argmax(dim=-1).item()
print(inference_model.config.id2label[prediction])
For a Hub adapter, pass its repository identifier to load_adapter(). For a local adapter, use its directory. Loading syntax can differ by release and by whether the head was saved, so test loading in a clean process rather than relying on the training process’s in-memory state.
Rank #4
Troubleshooting
Legacy imports appear in a tutorial
Examples using adapter-transformers or AutoModelWithHeads target the older ecosystem. Install adapters and load the model with AutoAdapterModel; do not mix forked legacy Transformers imports with current Adapters imports (AdapterHub documentation).
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
train_adapter() is missing
Check that the model came from adapters.AutoAdapterModel, not ordinary Transformers, and that you are not using a PEFT model. A quick check is:
print(type(model))
print(hasattr(model, "add_adapter"))
print(hasattr(model, "train_adapter"))
There is no active head or logits have the wrong shape
Set the correct active head, verify num_labels, preserve the expected label field, and ensure labels are valid integer IDs. Single-label binary classification is not the same as multilabel classification.
Loss does not improve
- Check label mapping and duplicate or leaked examples.
- Confirm that the adapter and head are active and trainable.
- Try overfitting a tiny subset to expose preprocessing or optimization errors.
- Review maximum length, class imbalance, learning rate, and domain mismatch.
CUDA out-of-memory errors occur
- Reduce batch size or maximum sequence length.
- Use gradient accumulation, mixed precision where supported, or gradient checkpointing.
- Keep dynamic padding enabled.
- Use a smaller RoBERTa checkpoint if necessary.
Adapters reduce trainable parameters and optimizer state; they do not remove the memory cost of loading the frozen base or storing activations.
The adapter will not load later
Confirm that the adapter directory contains weights and configuration, that the same compatible base model is used, and that the tokenizer, head, label mapping, and library versions match. Keep the base-model identifier alongside every adapter artifact.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBest Value
Task adapters, language adapters, and heads
A task adapter learns a downstream objective such as sentiment or intent classification. A language or domain adapter is generally learned from language-modeling data to improve representations and may later be composed with a task adapter; it is not automatically a classifier.
The adapter changes internal representations. The prediction head converts the final representation into class logits or regression values. For classification, the head is part of the usable task solution unless you have a separately managed compatible head.
Production checklist
- Save the exact base model and tokenizer identifiers.
- Record adapter type, configuration, bottleneck size, labels, and maximum length.
- Pin package versions and seeds; document hardware and precision settings.
- Keep validation and test data separate and report class-aware metrics.
- Check dataset privacy, licensing, and intended use before publishing.
- Publish a model card with provenance, limitations, evaluation results, and known failure cases.
- Monitor production drift and periodically reevaluate on representative labeled data.
When to use LoRA or full fine-tuning instead
Choose LoRA through PEFT when your organization already standardizes on PEFT or needs one workflow covering LoRA, IA3, AdaLoRA, and prefix tuning. Choose full fine-tuning when task-specific performance and flexibility outweigh larger artifacts, optimizer memory, and the loss of a shared frozen base. Neither choice is universally superior; compare methods on your dataset and deployment constraints.
Frequently Asked Questions
Can I use this adapter with RoBERTa-large?
Not automatically. Load the adapter with the exact compatible base architecture and record that identifier with the adapter; a RoBERTa-base adapter should not be assumed compatible with RoBERTa-large.
Does an adapter eliminate GPU memory problems?
No. It reduces trainable parameters and optimizer state, but the frozen base model and activation tensors still require memory.
Can I load a LoRA checkpoint with the Adapters library?
Not through the ordinary Adapters load_adapter() path. LoRA checkpoints use PEFT formats and APIs unless a documented integration explicitly supports that format.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

