René Cano
All projects

No. 9

FakeNewsDetector AI

News classifier with DistilBERT and heuristics

My role
Sole author
Period
Mar 2026

The problem

The hardest misinformation to catch isn't obvious clickbait but text that mimics the language of science: "researchers from a European university", "the study hasn't been published yet". A classifier trained only on obvious fake news tends to label that kind of text as reliable.

What I built

I built it alone and published it on March 30, 2026.

  • Data preparation: 5,000 articles per class from the ISOT dataset on Kaggle (Fake.csv and True.csv), with title and body joined and cut to 1,800 characters.
  • A DistilBERT fine-tune with the Hugging Face Trainer as a 2-class classifier, REAL and FAKE, with early stopping.
  • Heuristic rules: pseudoscience patterns (total protection, unpublished study, unnamed university), alarm signals and signals of a verifiable source.
  • Fusion of the model and the rules into 3 labels: Reliable, Doubtful and Fake. Doubtful shows up when the model can't separate the 2 classes well, and 2 or more pseudoscience patterns force Fake.
  • A Gradio interface that explains the verdict. Without a local model it uses public Hugging Face models and, as a last resort, the rules alone.

Architecture

Flow diagram: the ISOT dataset is prepared with 5,000 articles per class and trains a 2-class DistilBERT, REAL and FAKE; in the app, the news text goes through cleaning and detection of signals and pseudoscience patterns; the model output is fused with those rules, with public models or rules alone as fallback, and the Gradio interface shows Reliable, Doubtful or Fake.

Results

I don't publish an accuracy figure. The README reports a very high accuracy, but I don't consider it valid for two reasons I explain below: it's measured on the same set used to pick the model, and the dataset gives away the answer. An honest evaluation needs a separate, clean test set, and that's still pending.

Known limits

  • The README says the model is RoBERTa, but train.py fine-tunes DistilBERT. The README is wrong.
  • The evaluation set is the same one that drives early stopping and picks the best checkpoint, so the metric is optimistic.
  • In ISOT the real articles carry a "(Reuters)" dateline and the code doesn't strip it: the model can learn to recognize the news agency instead of truthfulness.
  • The trained model isn't in the repository, and the README mentions a notebook and a CSV of pseudoscience examples that aren't there either.
  • One of the fallback models is a sentiment analysis model (SST-2), not a news model.
  • The dataset is English only and there are no tests.

Links