All work

Fine-tuning study · February 2025

Teaching Llama 2 to answer questions about one university.

I wrote 41 question–answer pairs by hand and fine-tuned a 7B model on them, to find out how small a dataset could be and still change a model's behaviour. The answer turned out to be more interesting than the assistant.

Role
Solo, self-directed
Stack
PyTorch · transformers · peft · trl
Hardware
One Colab T4, free tier
Status
Complete, superseded

The problem

Abdul Wali Khan University Mardan gets the same questions every admissions cycle: which programmes exist, how to apply, what the deadlines are. The information is public and written down. It is just spread across pages nobody wants to read.

I had been reading about parameter-efficient fine-tuning and wanted to try it on something concrete rather than a tutorial dataset. A university assistant was small enough to finish and specific enough that I could tell whether the output was right.

What I built

I hand-wrote a dataset of 41 examples. Each one has a system prompt fixing the model's role as an academic advisor, a student question, and the answer I wanted back:

{
  "System Prompt": "You are an academic advisor for Abdul Wali Khan
                    University Mardan (AWKUM). Provide university-specific
                    information.",
  "User Prompt":   "What undergraduate programs are offered at AWKUM?",
  "Model Answer":  "AWKUM offers undergraduate programs in fields such as
                    Computer Science, Business Administration, English,
                    Physics, and Mathematics."
}

The model is Llama-2-7b-chat loaded in 4-bit NF4 quantisation, so it fits in a free Colab T4. On top of that sits a LoRA adapter at rank 64. The base weights stay frozen and only the adapter trains. Training runs through trl's SFTTrainer for one epoch.

The numbers

These are the values from the run, not rounded or reconstructed afterwards:

Training configuration and outcome
Base modelNousResearch/Llama-2-7b-chat-hf
Quantisation4-bit NF4, float16 compute, no nested quant
AdapterLoRA at rank 64, alpha 16, dropout 0.1
Dataset41 hand-written pairs
Epochs1
Batch size4 per device
Learning rate2e-4
Optimiser steps11
Wall-clock time8.86 seconds
Final training loss3.76
Held-out evaluationNone

A final loss of 3.76 after eleven steps is the whole story. The run was too short and the dataset too small for the adapter to move far from its initialisation. Whatever the model says afterwards, it is mostly still saying what Llama 2 already said.

What went wrong

I had no way to tell if it worked. I split nothing off for evaluation, so my entire test was typing questions in and reading the answers. That feels like evaluation and is not. I could not have told you whether the fine-tune helped, did nothing, or made things worse.

41 examples cannot carry facts. I expected the model to learn the content of the answers. It did not, and on reflection it could not. A fine-tune at this scale adjusts tone, format and the shape of a response. Facts have to come from somewhere the model can read at inference time. Ask it about a programme that was not in those 41 pairs and it produces something confident and invented.

The dataset was the real work and I treated it as setup. Writing 41 good pairs took longer than everything else combined, and 41 is still far too few. I spent my attention on the training configuration, which was the part a tutorial could have given me.

Example runs

The notebook saved one generation after training. Below: a training-set gold answer, that saved generation on a near-set question, and what happens outside the set, without inventing a transcript the notebook never kept.

1 · Inside the training set

Gold answer from awkum_dataset_updated.json, pair 0 of 41:

Q: What undergraduate programs are offered at AWKUM?
A: AWKUM offers undergraduate programs in fields such as Computer Science,
   Business Administration, English, Physics, and Mathematics.

This is the target the adapter was trained toward, not a free-form generation. It is short, specific, and bounded by what I wrote into the dataset.

2 · Near the training set, saved notebook generation

Prompt used in the notebook (with the original spelling): “How can i apply for addmission AWKUM?” The training set has a cleaner version of the same question. Truncated from the Colab cell output:

<s>[INST] How can i apply for addmission AWKUM? [/INST]
To apply for admission at Abdul Wali Khan University Mardan (AWKUM), you can
follow these steps:
 everybody can apply for admission at Abdul Wali Khan University Mardan
(AWKUM) by following these steps:

1. Visit the AWKUM website: Go to the official website of AWKUM at
   www.awkum.edu.pk
2. Check the admission criteria...
3. Download the application form...
4. Fill out the application form...
5. Pay the application fee...
6. Submit the application...
7. Appear in the entrance test (if required)...
8. Wait for the admission decision...

Compared with the gold answer (“apply through the online admission portal…”), this is longer, repeats itself, and fills in process detail that was never in the 41 pairs. Tone is university-ish; the facts are not audited.

3 · Outside the training set

The notebook only keeps that one generation call. I did not save an out-of-set transcript. When I typed questions about programmes, fees or deadlines that were not among the 41 pairs, the model still answered in a confident advisor voice, and invented details. That is the failure mode this project is meant to show: fluency without a source.

What I would do differently

I would not fine-tune at all. The task is retrieval: put the prospectus and admissions pages into a vector store, retrieve the relevant passages at query time, and let the model phrase an answer from text it can actually see. Facts stay in the documents, where they can be corrected without retraining anything.

I would also write the evaluation set before the model. Thirty questions with known answers, including questions the system should refuse, scored before and after any change. That is the piece missing from this project and from most of my earlier work, and it is what I am fixing now.

That rebuild is currently on my desk.