MMuflih Imaduddin
All projects

SBERT vs. GPT grading

Can a small fine-tuned model grade short answers as well as GPT?

Role
AI engineer · undergraduate thesis
Timeline
Apr 2026 – present
Stack
PythonOpenAISBERT
Links
Colab

Problem

Grading short answers by hand is slow; automated graders need to be accurate and affordable.

My role

Built the grading pipeline, an LLM-based data augmentation step and the benchmark.

Result

Compared zero-shot and few-shot approaches on F1-score, accuracy and API cost.

Overview

An Automated Short Answer Grading (ASAG) system that evaluates student answers with fine-tuned Sentence-BERT and GPT models, benchmarked side by side.

What I built

  • An LLM-driven augmentation pipeline that creates paraphrased and contradictory answers to balance the dataset.
  • Fine-tuned Sentence-BERT models and GPT prompting strategies (zero-shot and few-shot).
  • A benchmark on F1-score, accuracy and API inference cost, plus a demo app.

Screenshots

Next projectSIMAK Undip