~/posts/smollm-finetuning-comment-generator

4M scraped comments, 1.8M labels, one 3B model that writes in five styles

8 min read 1535 words
aillmlorapython

tl;dr

I deduplicated 4M+ scraped comments down to 3.7M, hand-labeled about 1k, grew that to about 8k with a DeBERTaV3 classifier and me correcting it, then auto-labeled the corpus and kept only confident labels (about 1.8M). A LoRA fine-tune of SmolLM3-3B on those writes comments I could only pick out 57% of the time in a 500-trial blind test.

I ended up with the kind of dataset that sounds like a flex and behaves like a responsibility: more than 4 million comments. Each one comes with the commenter’s username, a short description of what they were commenting on, and the comment itself.

I didn’t want a general chatbot out of it. I wanted something narrower: a model that writes believable comments in a style I pick. The styles I settled on:

  • happy
  • toxic
  • sarcastic
  • cringe
  • wholesome
  • noise: a special bucket for junk, left out of training later.

The scraping gets its own post. This one is about the part that actually hurt: cleaning, labeling, scaling the labels up, and turning all of it into a generator I can steer.

Cleaning comes first

Before labeling anything, I had to make sure I wasn’t labeling the same comment 50 times with different spellings. Scraped comments are full of:

  • exact duplicates (copy-paste, reposts, bots),
  • near duplicates (the same template with tiny edits),
  • empty rows and whitespace junk,
  • “lol”, “ok”, ”.”, and other short masterpieces.

The comments live in an SQLite database, so I wrote a cleanup pass that runs directly on it, in five steps.

1. Normalize. Collapse repeated whitespace, trim the edges, lowercase everything. “Nice!!!”, ” nice!!! ” and “NICE!!!” all become the same string, so the next steps can compare them.

2. Drop anything shorter than 12 characters. That removes a huge chunk of low-signal content: reactions, one-word replies, and fragments that only make sense inside their thread (and usually not even there).

3. Exact dedup. Each normalized comment gets a fast hash (xxHash), and a set in memory remembers which hashes were already seen. Seen before means duplicate, and it’s deleted. This removed 17k rows.

4. Near dedup. Exact matching isn’t enough, because the internet loves templates. Two comments count as duplicates when their Jaccard similarity is above a high threshold. Jaccard similarity is how much two sets overlap: shared items divided by all distinct items. Here the sets are the comments’ character n-grams, every run of a few consecutive characters.

A = "nice video!!"B = "nice video"nicicece·e·v·vivididedeoeo!o!!in both (8)only in A (2)Jaccard = shared / all distinct = 8 / 10 = 0.80
Illustrative example with 3-character n-grams (spaces shown as ·). Deleting the "!!" only removes two of ten n-grams, so the two comments score 0.80 and count as the same comment.

Comparing every comment with every other one is out of the question: 3.7 million comments make about 6.8 trillion pairs. MinHash shrinks each comment’s n-gram set into a short signature whose agreement with another signature estimates their Jaccard similarity. Locality-sensitive hashing (LSH) then puts similar signatures in the same buckets, so only comments that share a bucket get compared.

This catches the same sentence with its emojis removed, templates with a couple of words swapped, and tiny variations that are still the same comment. It removed 112k rows.

5. Work in chunks. Everything runs in batches, with batched reads and buffered bulk deletes, so it gets through millions of rows without turning my machine into a smoke test.

After cleaning, I had 3.7M comments, averaging 44 characters.

I also tracked how the distributions changed (length, emoji use, and so on) as a sanity check. The goal wasn’t to sterilize the dataset. It was to remove repetition and low-value sludge, so the models downstream would learn signal instead of copies.

Labels you can actually define

I kept the list of styles small on purpose. If you can’t define a label in one sentence, you don’t have a label, you have a future argument with yourself.

noise is a real label, not an afterthought. Spam, fragments with no context, unreadable junk, and anything I refuse to teach a model all go there. It never reaches the generator.

Each comment gets exactly one label (multiclass, not multi-label).

From 1k labels to 1.8M

Scraped comments: 4M+Scraped comments4M+After cleaning: 3.7MAfter cleaning3.7MConfident auto-labels: ~1.8MConfident auto-labels~1.8MReviewed labels: ~8kReviewed labels~8kHand labels: ~1kHand labels~1k
How much data made it through each stage. The bar in accent is what the generator was trained on. The labels I made or checked myself (about 1k, then about 8k) are barely visible at this scale.

1,000 labels by hand

I labeled about 1,000 comments myself. That was just enough to train a first classifier. Not great, not terrible, and enough to bootstrap the next step.

A first classifier

I fine-tuned DeBERTaV3 base as a classifier on the comment text alone. No tricks, no special losses. The goal was usefulness, not perfection.

It scored 85.22% accuracy on a validation set made of a random ~20% of my labels. I don’t treat that as a strong result. Random splits flatter a model when the data has near-duplicates, repeated templates, or topics that cluster, because similar comments land on both sides of the split. It was a “does this work at all?” signal while I iterated.

On a held-out split, the per-style F1 scores (a balance of how often a predicted label is right and how many of the true ones it finds, from 0 to 1) were:

StyleF1
happy0.84
toxic0.82
sarcastic0.80
cringe0.81
wholesome0.84

noise was tracked separately, since it’s excluded from training anyway. Good enough to bootstrap labeling, not good enough to trust blindly, so I filtered hard later.

8,000 labels with the model’s help

Then labeling became a loop, a form of active learning:

  1. Pick a comment.
  2. DeBERTa predicts its style.
  3. I accept or correct it.
  4. Every so often, retrain.
  5. Repeat.

I mixed three ways of doing this:

  • single random comments, accept or correct,
  • bulk review of whole batches,
  • once I had about 4k labels, generating comments in a target style, reviewing them, and moving the bad ones to noise.

That last one was surprisingly useful. It showed me how generation would fail long before I trained the generator, and it turned noise into a quality gate.

This got me to about 8k labels.

I also tried labeling with a bigger model, gpt-oss-safeguard-20b. It cost too much compute for this workflow. The small classifier plus my corrections was simpler and faster.

1.8M labels, filtered hard

Once the classifier was stable, I ran it over the whole cleaned corpus and kept only the confident answers:

  • keep a prediction only if its probability is at least 70%,
  • if noise gets more than 40%, drop the comment entirely,
  • leave noise out of the training set.

That left about 1.8M labeled comments. Fewer samples with clear labels beat millions of weak guesses that blur where one style ends and the next begins.

Training the generator

Base modelSmolLM3, 3B parameters
Trainingsupervised fine-tuning (SFT) only
MethodLoRA
Hardwareone RTX 4090
StackUnsloth, latest CUDA, TensorRT for inference

LoRA trains a small set of extra weights on top of the frozen base model instead of the whole model, which is what makes this fit on one consumer GPU. This was never going to replace an assistant. It writes believable comments in the style you ask for, and LoRA was the right tool for that.

Each training example is shaped like a chat:

  • System prompt: the style, plus a few guidelines.

  • User message:

    <username>...</username><description>...</description>
  • Assistant reply: the comment.

No transformations and no post-processing. Clean conditioning, direct generation.

Could I spot my own model?

I ran an arena-style blind test on myself:

  • The real comments came from the same distribution as the prompts, so the comparison was fair.
  • I was the only rater.
  • Each trial showed two comments, and I had to pick the one the model wrote.

Early quick checks used 100 trials. The latest run used 500, and I picked the model correctly about 57% of the time.

For this use case, that’s close enough to guessing to count as a win. It isn’t invisible, but it’s no longer obvious machine text. To be exact about it: 285 out of 500 is still better than a coin flip (a binomial test gives p ≈ 0.002, and the 95% interval is about 53% to 61%). I can spot it a little. Only a little.

What made it work

  1. Cleaning first. Dedup and dropping short comments kept the pipeline from learning repetition and noise.
  2. The classifier as a multiplier. 1k labels became 8k quickly once it could help.
  3. Confidence filtering. 4M raw comments are impressive. 1.8M confident labels are trainable.
  4. A strict noise bucket. Leaving garbage out is the difference between style control and style soup.
  5. LoRA on a small model. Cheap, practical, and good enough to be hard to spot.
hash: e0c
EOF