Rubric Grading - Tinker Documentation

Rubric-based Grading for LLMs

A simple example of using a grader LLM with rubrics

We show how to use a rubric-based LLM to provide a reward for an addition task. E.g.

**User**: What's 233 + 100?
**Assistant**: 333

Usually, this could be graded by matching the number to the ground truth 333 without needing an LLM. However, for pedagogical purposes, we will grade the response using a language model with a rubric. That is, we will ask a language model "Does the assistant answer 333?"

Generate an example dataset

To run this, first generate a dataset:

python -m tinker_cookbook.recipes.rubric.generate_data

Then you will see two jsonl files generated, one for training, one for testing. For example, if you look into tinker_cookbook/example_data/example_rubric_train.jsonl, each datapoint consists of

{
  "convo": [
    {
      "role": "user",
      "content": "What is 4 + 5?"
    },
    {
      "role": "assistant",
      "content": "9"
    },
    {
      "role": "user",
      "content": "What is 122 + 12?"
    }
  ],
  "rubric_items": [
    {
      "rubric_str": "Does the chatbot correctly get the answer 134?",
      "extraction_regex": "<score>(.*)</score>",
      "grader_output_format_instruction": "Please output your score between 0 and 1 wrapped in <score> ... </score>"
    }
  ]
}

Debugging and Printing What Happens During Rollouts

Run

python -m tinker_cookbook.recipes.rubric.debug_env

You can see the message that the policy sees, its response, the grader input, and the grader output.

An example training run

To train the LLM to add with a rubric-based LLM, run

python -m tinker_cookbook.recipes.rubric.train

You can see the reward quickly goes up. In this example, test/env/all/reward/total improves from 0.354 at step 0 to 0.994 by step 60, while the final training batch reaches env/all/rubric_score=1.0.

A more realistic dataset

We take the prometheus-eval/Feedback-Collection dataset from Hugging Face, which contains rubrics to grade general chat responses. Run the following to kick off training:

python -m tinker_cookbook.recipes.rubric.prometheus_experimental

We can see that the reward climbs up steadily.

Note that this training recipe is experimental -- to make the performance better we may need to fine-tune the grader LLM as well. We hope our code serves as a starting point for you to improve rubric-based grading for training LLMs!