Hugging Face's LLM Course: how the first three chapters take an FDE from pipeline() to fine-tuning
The course is free and available in many languages. Its first three chapters can be read as the order of work at a client site: set a baseline first, inspect the data next, fine-tune last.
In brief
- The course was written by ten authors from Hugging Face and its contributors. It is free, carries no ads, its code is licensed under Apache-2.0, and a Vietnamese translation is complete.
- The first three chapters follow the order an FDE works in at a client site: pipeline() for a baseline, tokenizers for understanding the data, Trainer for fine-tuning.
- Each chapter takes about 6-8 hours a week. At that pace, the first three chapters are enough to start a small project on real data of your own.
Picture a logistics company that wants to sort its customer-support tickets into five groups automatically. Nobody there asks what a transformer is. What they want to know is whether, within a few weeks, you can deliver a model that works on their own data.
The first three chapters of Hugging Face’s LLM Course come close to answering that question. Hugging Face does not frame the course as a project method, but read with an FDE’s eye, the order of the three chapters matches the sequence of work at a client site: build a baseline, inspect the data, and only then fine-tune.
The course is free, carries no ads and is open source under the Apache-2.0 licence. Its Vietnamese edition is listed among the completed translations. For a developer looking to move into forward deployed engineering, the cost of getting started is close to zero.
Who wrote it, and at what pace should you study?
The ten authors come from Hugging Face and its contributors, among them Sylvain Gugger, Lewis Tunstall, Leandro von Werra, Merve Noyan and Ben Burtenshaw. The official page is titled “Welcome to the 🤗 Course!”, and the citation reads “The Hugging Face Course, 2022”.
The GitHub README describes the goal as teaching how to apply Transformers to a range of natural language processing tasks and beyond. The suggested pace is one chapter a week, at roughly 6-8 hours a week.
The repo is open for anyone to read, but Hugging Face is not currently accepting community contributions for new chapters. For learners this changes nothing: you are there to study and read code, and the first three chapters are enough to get to work.
Idea one: build a baseline first, choose a model later
Chapter 1 teaches the pipeline() function for tasks such as text generation and classification. It sounds simple, but it is the habit that separates an FDE from someone who only builds demos.
Back to the logistics company. On day one, fine-tuning should not come up. Take a few dozen real tickets, run them through an off-the-shelf classification pipeline and count how many are misclassified. That number is your baseline: the mark every later proposal has to beat.
If the baseline is already good enough, you have just saved the client several weeks of work. If it is not, the list of misclassified tickets tells you where to look.
Idea two: the tokenizer is where real data gets distorted
Chapter 2 defines the tokenizer’s job concisely: turning text into data a model can process. It then compares three ways of splitting text: by word (word-based), by character (character-based) and by units smaller than a word (subword). The subword family includes BPE, used in GPT-2; WordPiece, used in BERT; and SentencePiece/Unigram.
This is the chapter to read most slowly, because client data is never as clean as the examples. Suppose the logistics tickets contain lines like “ko đc giao hàng”, Vietnamese shorthand for “didn’t get delivered”, typed without diacritics, or have a waybill number pasted mid-sentence.
In principle, a word-based tokenizer will often meet words outside its vocabulary. A subword tokenizer breaks unfamiliar words into pieces it already knows, but whether those pieces keep their meaning is something you only find out by looking.
The first thing to do is run the tokenizer of the model you intend to use on those very tickets, and look at the splits with your own eyes. Five minutes of this can explain why the baseline fails, before anyone gets round to blaming the model.
Idea three: fine-tuning is mostly data preparation
Chapter 3 teaches fine-tuning with the high-level Trainer API, following modern best practices. But the other half of the chapter, preparing a large dataset from the Hub with the new features of 🤗 Datasets, is the part closest to real work.
In the logistics example, the hard part is not calling Trainer. It is turning a few thousand messy tickets into a dataset with clear labels, a sensible train/test split, and the same tokenizer you checked in chapter 2.
Once that is done, all that remains is to compare the fine-tuned model with the baseline from chapter 1. That comparison is what you present to the client.
Who should take it, and in what order?
Go strictly in order, 1, 2, 3, one chapter a week.
The reason is practical: you only know whether a fine-tuned model was worth the effort once you have the baseline from chapter 1, and the dataset in chapter 3 has to pass through the tokenizer you came to understand in chapter 2.
If a translation exists in your language, it helps, but use it alongside the English original. Read the translation to grasp the concepts, and read code and API names in English, because documentation, issues and, later, client questions all use English terminology.
When reading job descriptions, watch for terms such as Transformers, fine-tuning, tokenization and Hugging Face. On your CV, do not write “completed the course”. Describe a three-step project instead: a baseline with pipeline, a tokenizer analysis on real data, and a fine-tuned model that beats the baseline, with measured numbers and a link to the repo.
A certificate shows that you studied. A project done in the order of these three chapters shows that you know where to start when you arrive at a client.