LLM Fine-Tuning with GRPO (RL)
An AI-guided learning project: end-to-end reinforcement-learning fine-tuning of Llama 3.2 3B Instruct with TRL's GRPOTrainer. I applied 4-bit NF4 quantization with LoRA adapters (r=16) to fit training in 16 GB VRAM, designed a rule-based reward function over a 60-prompt custom dataset, tracked runs with W&B, and published the adapter to Hugging Face with an honest model card.
View on Hugging Face →