Multi-Domain GRPO Reasoning
Google Tunix Hackathon project teaching Gemma2 to reason transparently with multi-domain GRPO training.
Multi-Domain GRPO Reasoning explores transparent reasoning training for Gemma2 across multiple domains.
- Used GRPO training to improve reasoning behavior and formatting discipline.
- Reached 100% format accuracy and 60.8% exact accuracy.
- Built for the Google Tunix Hackathon using Gemma, JAX, and TPU tooling.