Multi-Domain GRPO Reasoning

Google Tunix Hackathon project teaching Gemma2 to reason transparently with multi-domain GRPO training.

Multi-Domain GRPO Reasoning explores transparent reasoning training for Gemma2 across multiple domains.

  • Used GRPO training to improve reasoning behavior and formatting discipline.
  • Reached 100% format accuracy and 60.8% exact accuracy.
  • Built for the Google Tunix Hackathon using Gemma, JAX, and TPU tooling.

View the repository