# Hugging Face's LLM Course: how the first three chapters take an FDE from pipeline() to fine-tuning

> The course is free and available in many languages. Its first three chapters can be read as the order of work at a client site: set a baseline first, inspect the data next, fine-tune last.

Bản gốc: https://fdetimes.net/en/books-courses/hugging-face-llm-course-first-three-chapters/

Picture a logistics company that wants to sort its customer-support tickets into five groups automatically. Nobody there asks what a transformer is. What they want to know is whether, within a few weeks, you can deliver a model that works on their own data.

The first three chapters of Hugging Face's LLM Course come close to answering that question. Hugging Face does not frame the course as a project method, but read with an FDE's eye, the order of the three chapters matches the sequence of work at a client site: build a baseline, inspect the data, and only then fine-tune.

The course is free, carries no ads and is open source under the Apache-2.0 licence. Its Vietnamese edition is listed among the completed translations. For a developer looking to move into forward deployed engineering, the cost of getting started is close to zero.

## Who wrote it, and at what pace should you study?

The ten authors come from Hugging Face and its contributors, among them Sylvain Gugger, Lewis Tunstall, Leandro von Werra, Merve Noyan and Ben Burtenshaw. The official page is titled "Welcome to the 🤗 Course!", and the citation reads "The Hugging Face Course, 2022".

The GitHub README describes the goal as teaching how to apply Transformers to a range of natural language processing tasks and beyond. The suggested pace is one chapter a week, at roughly 6-8 hours a week.

The repo is open for anyone to read, but Hugging Face is not currently accepting community contributions for new chapters. For learners this changes nothing: you are there to study and read code, and the first three chapters are enough to get to work.

## Idea one: build a baseline first, choose a model later

Chapter 1 teaches the `pipeline()` function for tasks such as text generation and classification. It sounds simple, but it is the habit that separates an FDE from someone who only builds demos.

Back to the logistics company. On day one, fine-tuning should not come up. Take a few dozen real tickets, run them through an off-the-shelf classification pipeline and count how many are misclassified. That number is your baseline: the mark every later proposal has to beat.

If the baseline is already good enough, you have just saved the client several weeks of work. If it is not, the list of misclassified tickets tells you where to look.

## Idea two: the tokenizer is where real data gets distorted

Chapter 2 defines the tokenizer's job concisely: turning text into data a model can process. It then compares three ways of splitting text: by word (word-based), by character (character-based) and by units smaller than a word (subword). The subword family includes BPE, used in GPT-2; WordPiece, used in BERT; and SentencePiece/Unigram.

This is the chapter to read most slowly, because client data is never as clean as the examples. Suppose the logistics tickets contain lines like "ko đc giao hàng", Vietnamese shorthand for "didn't get delivered", typed without diacritics, or have a waybill number pasted mid-sentence.

In principle, a word-based tokenizer will often meet words outside its vocabulary. A subword tokenizer breaks unfamiliar words into pieces it already knows, but whether those pieces keep their meaning is something you only find out by looking.

The first thing to do is run the tokenizer of the model you intend to use on those very tickets, and look at the splits with your own eyes. Five minutes of this can explain why the baseline fails, before anyone gets round to blaming the model.

**Điểm mấu chốt:** Before blaming the model, look at how the tokenizer has cut up the client's data.

## Idea three: fine-tuning is mostly data preparation

Chapter 3 teaches fine-tuning with the high-level Trainer API, following modern best practices. But the other half of the chapter, preparing a large dataset from the Hub with the new features of 🤗 Datasets, is the part closest to real work.

In the logistics example, the hard part is not calling `Trainer`. It is turning a few thousand messy tickets into a dataset with clear labels, a sensible train/test split, and the same tokenizer you checked in chapter 2.

Once that is done, all that remains is to compare the fine-tuned model with the baseline from chapter 1. That comparison is what you present to the client.

## Who should take it, and in what order?

Go strictly in order, 1, 2, 3, one chapter a week.

The reason is practical: you only know whether a fine-tuned model was worth the effort once you have the baseline from chapter 1, and the dataset in chapter 3 has to pass through the tokenizer you came to understand in chapter 2.

If a translation exists in your language, it helps, but use it alongside the English original. Read the translation to grasp the concepts, and read code and API names in English, because documentation, issues and, later, client questions all use English terminology.

When reading job descriptions, watch for terms such as Transformers, fine-tuning, tokenization and Hugging Face. On your CV, do not write "completed the course". Describe a three-step project instead: a baseline with pipeline, a tokenizer analysis on real data, and a fine-tuned model that beats the baseline, with measured numbers and a link to the repo.

A certificate shows that you studied. A project done in the order of these three chapters shows that you know where to start when you arrive at a client.

**Thử ngay tuần này:**

- Work through chapter 1, then use pipeline() to classify 20 real pieces of text from your own work. Note the ones the model gets wrong.
- Run two different tokenizers on the same 20 texts and compare how each handles abbreviations and words written without accents or diacritics.
- Open the huggingface/course repo on GitHub, find the edition in your language, and schedule chapters 1-3 over the next three weeks.

## Nguồn

- [Welcome to the 🤗 Course! (Hugging Face LLM Course)](https://huggingface.co/learn/llm-course/chapter1/1)

- [Tokenizers - Hugging Face LLM Course](https://huggingface.co/learn/llm-course/chapter2/4)

- [Introduction - Hugging Face LLM Course (Chapter 3)](https://huggingface.co/learn/llm-course/chapter3/1)

- [huggingface/course: The Hugging Face course on Transformers (GitHub)](https://github.com/huggingface/course)
