TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Fine-Tune a Local LLM With QLoRA and Hugging Face PEFT
Learning Hub

How to Fine-Tune a Local LLM With QLoRA and Hugging Face PEFT

Learn how to fine-tune a small local LLM with QLoRA and Hugging Face PEFT, teaching it a strict response format on a single consumer GPU with under 3GB of VRAM.

August 7, 2026 15 Min Read
96

Fine-tuning a language model sounds like something only a well-funded lab can do. In practice, you can teach a small open model a new habit, a fixed response format, a house style, a narrow skill, on a single consumer GPU with less VRAM than a modern game uses, in under two minutes of actual training time. This tutorial builds that end to end: you will take Qwen2.5-1.5B-Instruct, a small open-weight chat model, and use QLoRA to teach it to always answer operations questions in a strict three-line DIAGNOSIS, FIX, VERIFY format instead of free-form prose, then prove the change is real by switching the adaptation on and off on the exact same loaded model.

Table Of Content

  • What You Will Build
  • Prerequisites
  • Understanding the Concepts First
  • What Supervised Fine-Tuning Actually Changes
  • What LoRA Changes About Training
  • What QLoRA Adds: Fitting It All in Less VRAM
  • Step 1: Set Up an Isolated Python Environment
  • Step 2: Confirm Your GPU Is Visible to PyTorch
  • Step 3: Install the Fine-Tuning Stack
  • Step 4: Build a Small, Focused Training Dataset
  • Step 5: Write the Training Script
  • Loading the Base Model in 4-Bit
  • Formatting the Dataset With the Model’s Own Chat Template
  • Attaching a LoRA Adapter
  • Configuring and Running SFTTrainer
  • Step 6: Run the Fine-Tune and Read the Results
  • Step 7: Compare the Base Model to the Fine-Tuned Model
  • Common Mistakes and Gotchas
  • How to Confirm It All Worked End to End
  • Next Steps

This matters for anyone building an AI agent or internal tool: prompting a model to follow a format works most of the time, but a fine-tuned model follows it far more reliably, because the behavior is baked into its weights instead of relying on the model remembering an instruction buried in a long system prompt. That reliability is exactly what a tool-calling agent or a structured-output pipeline needs.

Before writing any code, here is what the key terms mean, since the rest of this tutorial builds directly on them. Fine-tuning means taking a model that is already trained and training it further on your own examples so its behavior shifts toward what you want. Supervised fine-tuning, or SFT, is the specific case where those examples are labeled input and output pairs, in this case a question and the exact answer format you want back. LoRA, short for Low-Rank Adaptation, is a technique that fine-tunes a model without touching most of its original weights. QLoRA combines LoRA with 4-bit quantization so the whole process fits in a fraction of the GPU memory a full fine-tune would need. Each of these gets a fuller explanation before you use it below.

What You Will Build

Working in a single Python project, you will:

  • Load Qwen2.5-1.5B-Instruct in 4-bit precision with bitsandbytes.
  • Attach a LoRA adapter with Hugging Face PEFT and train it on 20 handwritten ops-triage examples using TRL’s SFTTrainer.
  • Watch the real training loss and token accuracy move as the model learns the format.
  • Compare the base model’s answers against the fine-tuned model’s answers on three questions it never saw during training.
  • Prove the adapter, not some other change, is responsible for the new behavior by disabling it on the same loaded model.

Prerequisites

  • Python 3.11 or later, with pip available. This tutorial was built and tested on Python 3.13.14.
  • An NVIDIA GPU with CUDA support is strongly recommended. This tutorial was tested on an NVIDIA RTX A400 with 4GB of VRAM, and peak usage during training stayed under 2.8GB. If you only have a CPU, the same code still runs; training will just take much longer, since the source article this tutorial was inspired by was itself written and tested on a CPU-only MacBook Pro.
  • About 4GB of free disk space for the base model download.
  • Comfort with the command line and basic Python. No prior fine-tuning experience is required; every concept is defined before it is used.
  • Commands below use Windows PowerShell / Git Bash syntax for activating a virtual environment; substitute source venv/bin/activate on macOS or Linux.

Understanding the Concepts First

What Supervised Fine-Tuning Actually Changes

A pretrained chat model has already learned language, facts, and a general sense of how to hold a conversation from a massive, generic training run. Supervised fine-tuning does not repeat that process. It shows the model a smaller, focused set of question-and-answer examples and adjusts its weights so it becomes more likely to produce that kind of answer for that kind of question. With only a handful of examples, SFT is far better at teaching a model a format, a tone, or a narrow behavior than at teaching it genuinely new facts. You will see that distinction play out directly in this tutorial’s own results.

What LoRA Changes About Training

A full fine-tune updates every one of a model’s weights, which for even a small 1.5-billion-parameter model means gradients and optimizer state for 1.5 billion numbers. Hugging Face’s PEFT library, short for Parameter-Efficient Fine-Tuning, takes a different approach. Its LoRA method freezes the entire original model and instead represents the weight updates as two much smaller matrices, injected alongside the frozen weights it is adapting. The original weights never change; only these small added matrices are trained, and the two are combined at inference time to produce the final result. Two settings control this: r, the rank, controls how large those added matrices are and therefore how many parameters actually get trained, and lora_alpha is a scaling factor that controls how strongly the adapter’s output influences the frozen model’s output. LoRA is normally applied only to a model’s attention layers rather than every layer, which is why you will see target_modules set to the model’s query, key, value, and output projections below.

Because only those small matrices are trained and saved, a LoRA adapter file ends up tiny compared to the base model, and you can swap different adapters in and out of the same frozen base model without ever touching it. You will see exactly how tiny later in this tutorial, with real file sizes from this exact run.

What QLoRA Adds: Fitting It All in Less VRAM

LoRA alone still requires loading the full base model at 16-bit precision, which for larger models is more GPU memory than most people have. The QLoRA paper (Dettmers et al., 2023) solved this by quantizing, meaning compressing, the frozen base model’s weights down to 4 bits each before training the LoRA adapter on top of them. The paper reports this makes it possible to “finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance.” QLoRA introduces a 4-bit format called NF4, or 4-bit NormalFloat, which the bitsandbytes documentation describes as “a quantization data type for normally distributed data” that improves accuracy over a plain 4-bit float for typical model weights. It also adds double quantization, which quantizes the small quantization constants themselves to save a little more memory, and paged optimizers, which manage memory spikes during training. Put together, these are the exact settings you will pass to BitsAndBytesConfig in Step 5.

Step 1: Set Up an Isolated Python Environment

Keep this project in its own virtual environment so its dependencies cannot collide with anything else on your machine.

mkdir qlora-tutorial
cd qlora-tutorial
python -m venv venv
venv\Scripts\activate
python -m pip install --upgrade pip

On macOS or Linux, activate with source venv/bin/activate instead. You will know it worked because your shell prompt now shows (venv) at the start of the line.

Step 2: Confirm Your GPU Is Visible to PyTorch

Install PyTorch first, using the CUDA 12.4 wheel index, then check that it can actually see your GPU before installing anything else. It is much easier to debug a GPU problem now than after you have five more libraries layered on top.

pip install torch --index-url https://download.pytorch.org/whl/cu124
python -c "import torch; print(torch.__version__); print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"

On the machine used for this tutorial, that printed:

2.6.0+cu124
True
NVIDIA RTX A400

A common point of confusion here: running nvidia-smi may show a much newer CUDA version, for example CUDA 13.2, than the cu124 wheel you just installed. That is fine. The number in nvidia-smi is the newest CUDA version your driver supports, not a version you must match exactly; PyTorch’s bundled CUDA 12.4 runtime works on any driver that supports CUDA 12.4 or newer. If torch.cuda.is_available() prints False, your driver is too old, no compatible GPU is present, or you accidentally installed the CPU-only build; reinstall with the command above and double check nvidia-smi runs at all first.

Step 3: Install the Fine-Tuning Stack

Install the remaining libraries. This tutorial pins the exact versions actually used, so you can reproduce the same results:

pip install transformers datasets trl peft accelerate bitsandbytes numpy

Versions used and verified for this tutorial: transformers 5.14.1, datasets 5.0.1, trl 1.9.2, peft 0.20.0, accelerate 1.14.0, bitsandbytes 0.50.0. TRL (Transformers Reinforcement Learning) is the Hugging Face library that provides SFTTrainer, the training loop you will use in Step 5.

One tool intentionally left out: this tutorial does not use Unsloth, a popular fine-tuning speed-up library, even though the article that inspired this one does. Unsloth was tried here too, and on this native Windows setup it failed on the very first import, deep inside its own device-detection code, with NotImplementedError: Unsloth cannot find any torch accelerator? You need a GPU, despite torch.cuda.is_available() already confirming the GPU was working correctly one step earlier. Rather than debug a third-party library’s device detection, this tutorial uses the standard Transformers plus PEFT plus TRL stack instead, which hit no such error and implements the same underlying QLoRA algorithm; you may get faster training with Unsloth on a platform where it installs cleanly for you, not different correctness.

Step 4: Build a Small, Focused Training Dataset

SFT data is just a list of question-and-answer pairs, formatted the way you want the model to actually respond. This tutorial uses 20 short, handwritten operations-troubleshooting questions, each answered in a strict three-field format: a one-line DIAGNOSIS, a concrete FIX with a real command, and a VERIFY step to confirm the fix worked. Save this as train_data.json; a few representative entries look like this:

[
  {
    "question": "My systemd service keeps restarting every few seconds, how do I figure out why?",
    "answer": "DIAGNOSIS: The service is crashing right after startup, so systemd's restart policy keeps retrying it.\nFIX: Run systemctl status myservice and journalctl -u myservice -n 50 --no-pager to read the crash reason in the logs.\nVERIFY: After fixing the root cause, run systemctl restart myservice then systemctl is-active myservice and confirm it prints active."
  },
  {
    "question": "nginx returns a 502 Bad Gateway.",
    "answer": "DIAGNOSIS: nginx could not get a valid response from the upstream application, so the backend is likely down or unreachable.\nFIX: Check the upstream process with systemctl status myapp, and confirm proxy_pass in the nginx config matches the app's real listening port.\nVERIFY: Run curl -I http://localhost:<port> directly against the upstream, then reload the page and confirm it returns 200."
  }
]

The full set covers 20 different, unrelated failure scenarios (disk space, DNS, Docker, Kubernetes, SSH, cron, TLS certificates, and more) so the model learns the shape of a good answer rather than memorizing one topic. Twenty examples is deliberately small. That is enough to teach a strict, mechanical format, which is what this tutorial demonstrates, but it is not enough to teach the model new facts it did not already know; the Common Mistakes section below shows exactly where that distinction shows up in real output.

Step 5: Write the Training Script

Create train.py. This section walks through it in four pieces: loading the model in 4-bit, formatting the dataset, attaching the LoRA adapter, and running the trainer.

Loading the Base Model in 4-Bit

import json
import torch
from datasets import Dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from trl import SFTTrainer, SFTConfig

MODEL_ID = "Qwen/Qwen2.5-1.5B-Instruct"
OUTPUT_DIR = "./triage-lora-adapter"
SYSTEM_PROMPT = "You are an ops triage assistant. Always answer in exactly three lines: DIAGNOSIS, FIX, VERIFY."

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    quantization_config=bnb_config,
    device_map={"": 0},
)

The four bnb_4bit_* settings are the QLoRA techniques from the concepts section, made concrete: load_in_4bit and bnb_4bit_quant_type="nf4" load and quantize the frozen base weights to the NF4 format, bnb_4bit_use_double_quant turns on double quantization for extra memory savings, and bnb_4bit_compute_dtype tells the GPU to do the actual matrix math in bfloat16 even though the weights are stored in 4 bits. device_map={"": 0} pins the whole model to GPU 0 rather than letting it split across devices.

Formatting the Dataset With the Model’s Own Chat Template

with open("train_data.json") as f:
    examples = json.load(f)

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

def to_text(example):
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": example["question"]},
        {"role": "assistant", "content": example["answer"]},
    ]
    return {"text": tokenizer.apply_chat_template(messages, tokenize=False)}

dataset = Dataset.from_list(examples).map(to_text)

Every model has its own chat template, the exact special tokens and layout it was trained to expect around system, user, and assistant turns. apply_chat_template builds that layout for you instead of you guessing at Qwen’s specific format, which matters because training on the wrong template teaches the model to associate your answers with a format it has never actually seen.

Attaching a LoRA Adapter

model = prepare_model_for_kbit_training(model)

lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()

prepare_model_for_kbit_training is not optional here. Per Hugging Face’s own PEFT quantization guide, a quantized model is not typically trained further on its own, since training directly on lower-precision weights and activations can be unstable; PEFT adapters get around that by only adding small, full-precision trainable parameters on top, and this function preprocesses the quantized model so that addition works correctly. target_modules lists Qwen’s four attention projections, matching PEFT’s own documented example for QLoRA-style training. On this run, print_trainable_parameters() reported:

trainable params: 4,358,144 || all params: 1,548,072,448 || trainable%: 0.2815

Under a third of one percent of the model’s parameters are actually being trained. Everything else stays frozen and quantized.

Configuring and Running SFTTrainer

sft_config = SFTConfig(
    output_dir=OUTPUT_DIR,
    num_train_epochs=12,
    per_device_train_batch_size=2,
    learning_rate=2e-4,
    logging_steps=5,
    save_strategy="no",
    report_to="none",
    max_length=512,
    bf16=True,
    dataset_text_field="text",
    seed=42,
)

trainer = SFTTrainer(model=model, args=sft_config, train_dataset=dataset)
train_result = trainer.train()

trainer.save_model(OUTPUT_DIR)
tokenizer.save_pretrained(OUTPUT_DIR)

With only 20 examples, 12 epochs (12 full passes over the data, 120 total training steps at batch size 2) gives the model enough repetition to firmly learn the format without an excessive run time. save_strategy="no" skips writing checkpoints during training since this run is short enough not to need one; trainer.save_model at the end writes the final adapter.

Step 6: Run the Fine-Tune and Read the Results

python train.py

On the RTX A400 used for this tutorial, the base model loaded in 41.9 seconds and used 1.15GB of VRAM just sitting there in 4-bit. Training then ran for 115.7 seconds and peaked at 2.74GB of VRAM, comfortably inside a 4GB card. Here is the real, unedited start and end of the logged training loss:

{'loss': '2.935', ... 'mean_token_accuracy': '0.4889', 'epoch': '0.5'}
{'loss': '2.169', ... 'mean_token_accuracy': '0.5549', 'epoch': '1'}
{'loss': '1.216', ... 'mean_token_accuracy': '0.7163', 'epoch': '2.5'}
...
{'loss': '0.1999', ... 'mean_token_accuracy': '0.9536', 'epoch': '9.5'}
{'loss': '0.118',  ... 'mean_token_accuracy': '0.9755', 'epoch': '11.5'}
{'loss': '0.1654', ... 'mean_token_accuracy': '0.9629', 'epoch': '12'}

Loss is a single number representing how surprised the model is by the correct next token; lower means the model is predicting the training answers more confidently. mean_token_accuracy is more intuitive: the fraction of tokens where the model’s top guess matched the actual training answer. It climbed from 49 percent right after the first half-epoch to over 96 percent by the end, which is exactly what you want to see: steady, consistent improvement rather than a flat line (which would mean the model is not learning) or a sudden collapse to zero (which usually means the learning rate is too high).

Step 7: Compare the Base Model to the Fine-Tuned Model

Create inference.py to load the base model once, generate answers with no adapter, then attach the adapter and generate again on the same three questions, none of which appeared anywhere in the 20 training examples:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import PeftModel

MODEL_ID = "Qwen/Qwen2.5-1.5B-Instruct"
ADAPTER_DIR = "./triage-lora-adapter"
SYSTEM_PROMPT = "You are an ops triage assistant. Always answer in exactly three lines: DIAGNOSIS, FIX, VERIFY."

TEST_QUESTIONS = [
    "My Redis server keeps evicting keys unexpectedly.",
    "kubectl apply says the deployment is stuck at 0/3 ready replicas.",
    "My cron-scheduled backup script silently stopped producing new backup files.",
]

def generate(model, tokenizer, question):
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": question},
    ]
    # apply_chat_template with tokenize=True returns a BatchEncoding (a dict
    # with input_ids and attention_mask), not a raw tensor, on transformers
    # 5.14.1. Unpack it with ** so generate() receives both fields; passing
    # it directly as input_ids= raises AttributeError deep inside generate().
    inputs = tokenizer.apply_chat_template(
        messages, tokenize=True, add_generation_prompt=True, return_tensors="pt"
    ).to(model.device)
    outputs = model.generate(**inputs, max_new_tokens=120, temperature=0.2, do_sample=True)
    generated = outputs[0][inputs["input_ids"].shape[-1]:]
    return tokenizer.decode(generated, skip_special_tokens=True)

That comment documents a real bug hit while building this tutorial, covered in full in the next section. With that resolved, loading the base model, generating baseline answers, then wrapping it with PeftModel.from_pretrained(base_model, ADAPTER_DIR) and generating again produced this, captured verbatim:

BASELINE (no adapter)

Q: My Redis server keeps evicting keys unexpectedly.
The issue is likely due to a misconfiguration or race conditions in your
Redis configuration settings. Ensure that the eviction policy and threshold
values are correctly set according to your needs. To fix this, review your
Redis configuration file (usually redis.conf) and adjust the eviction
policies as necessary...

FINE-TUNED (LoRA adapter active)

Q: My Redis server keeps evicting keys unexpectedly.
DIAGNOSIS: Eviction is likely when a key's expiration time has passed but
no callback was registered to keep the entry alive.
FIX: Add a callback with expireafter_seconds set on the key, then remove
the callback before expiry to prevent it from being evicted prematurely.
VERIFY: Retry fetching the key and confirm it returns the expected value
without retrying indefinitely.

Q: kubectl apply says the deployment is stuck at 0/3 ready replicas.
DIAGNOSIS: The pod status shows only zero out of three pods are ready.
FIX: Run kubectl get pod -n <deployment> and confirm all containers have
Ready=true, then run kubectl describe pod <selector> for details.
VERIFY: Run kubectl get deploy -n <deployment> and confirm the replica
count is correct and all conditions are resolved.

The baseline model answers in reasonable but unstructured prose. The fine-tuned model, on questions it never trained on, reliably snaps into the exact three-line format for every single test question. That is the format transfer working.

To prove the adapter itself is responsible, and not some other change, add one more block that regenerates the first question with the adapter temporarily switched off on the very same loaded model:

with ft_model.disable_adapter():
    print(generate(ft_model, tokenizer, TEST_QUESTIONS[0]))

The output with the adapter disabled came back byte-for-byte identical to the original baseline paragraph above. Same weights, same question, same random seed, adapter off, prose again. That is direct proof the LoRA weights, and nothing else, are causing the behavior change, and that they can be toggled on and off without ever modifying the underlying base model.

Common Mistakes and Gotchas

  • Treating apply_chat_template‘s output as a raw tensor. On transformers 5.14.1, calling it with tokenize=True, return_tensors="pt" returns a BatchEncoding, a dict-like object with input_ids and attention_mask keys, not a tensor. Passing that object straight in as input_ids=inputs compiles and runs right up until generate() tries to read inputs_tensor.shape[0] internally, then fails with a confusing AttributeError. Unpack it with **inputs instead, as shown in Step 7.
  • Skipping prepare_model_for_kbit_training. It is easy to assume you can quantize a model and immediately wrap it in get_peft_model. Hugging Face’s own quantization guide is explicit that a quantized model needs this preprocessing step first; skipping it risks unstable training on top of low-precision weights.
  • Expecting fine-tuning to teach new facts from 20 examples. Look closely at the Redis answer above: it invents a nonexistent expireafter_seconds callback mechanism. Redis has no such feature. The model learned the DIAGNOSIS/FIX/VERIFY format perfectly, but with only 20 tiny examples and no Redis-specific training data, it filled in plausible-sounding technical details it was never taught, a textbook example of hallucination. SFT with LoRA on a small dataset is reliable for teaching style, tone, and format; it is not a substitute for giving the model accurate information it does not already have. For that, pair fine-tuning with retrieval instead, covered in Next Steps below.
  • Running out of VRAM on a bigger model. This tutorial’s 1.5B model peaked at 2.74GB during training. Larger base models need proportionally more, even quantized to 4-bit. If you hit an out-of-memory error on a bigger model, lower per_device_train_batch_size to 1 before reaching for a smaller model.
  • Forgetting to set a pad token. Qwen2.5’s tokenizer has no default pad_token. Training will error out on batching without the if tokenizer.pad_token is None: tokenizer.pad_token = tokenizer.eos_token line from Step 5.

How to Confirm It All Worked End to End

  • Training loss should trend steadily downward across the run, not stay flat or spike to nan. This run went from 2.935 to 0.1654.
  • mean_token_accuracy should climb toward, and likely past, 90 percent by the final epochs on a small, repetitive dataset like this one. This run reached 96 percent.
  • On questions the model never saw during training, the fine-tuned model should reliably follow your target format, while the base model does not.
  • Disabling the adapter with model.disable_adapter() on the same loaded model should reproduce the original base behavior exactly, confirming the adapter, and only the adapter, changed anything.
  • Check the adapter’s file size. In this run, adapter_model.safetensors was 8,746,152 bytes, about 8.7MB, versus 3,087,467,144 bytes, about 3.09GB, for the base model’s own weight file: the adapter is roughly 350 times smaller than the model it modifies. That is the practical payoff of LoRA: you can store, version, and swap many small adapters without ever duplicating the base model.

Next Steps

  • Swap in your own dataset. The entire technique here works identically for any narrow behavior you can write 15 to 30 clear examples of: a support tone, a structured extraction format, a domain-specific style.
  • For deployment, call model.merge_and_unload() on the PEFT model to bake the adapter weights permanently into a standalone copy of the base model, so you can serve it without the PEFT library at inference time.
  • If your use case needs the model to know specific facts, documents, or up-to-date information rather than just follow a format, pair this technique with retrieval instead of trying to train facts in. See this site’s tutorial on building a local RAG Q&A agent with LangChain and Ollama for that complementary approach.
  • Once you are comfortable with the basics here, experiment with a higher LoRA rank (r) on a larger, more diverse dataset, and watch how training time and VRAM use scale.

Tags:

huggingface-peftloraMachine LearningPythonqlora

Share

A curved steel highway guardrail running along a rocky coastline
Previous Post

AWS Turns Bedrock Guardrail Violations Into Standard Security Telemetry

Several colorful kites fly over a beach in De Haan, Belgium, as kitesurfers ride the waves below
Next Post

Cloudflare Launches Kitesurf, a Stripped-Down Browser for AI Agents

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026