# Session Metrics

Session metrics provide critical insight into your session's activity, allowing you to monitor usage and debug performance bottlenecks. The **Session Metrics** view shows exactly what your training and sampling traffic did over time, including throughput, training utilization, and in-flight requests. You can read the summary panels in the console to track general resource utilization, or export the full timeline to [Perfetto](https://ui.perfetto.dev/) for a detailed, per-request view.

_The Session Metrics view in the console._

The console panels are a subset of the trace: **In-flight sample requests** (the console's "Active sample requests"), **Training utilization**, and the **Training**, **Generation**, and **Prefill tokens/sec** rates. The exported Perfetto trace is a strict superset, adding per-request sizes, latency percentiles, totals, rate-limit events, and the individual request spans.

## Exporting and viewing the trace

For the full picture, every request and counter at fine resolution, export a Perfetto trace:

1. Click **Download Perfetto** on the session page (large sessions take a moment to build).
2. Open **[ui.perfetto.dev](https://ui.perfetto.dev/)**.
3. **Drag the `.pftrace` in**, or use **Open trace file**.

Counter tracks run along the top (grouped: Concurrency, Throughput, Per Request, Latency, Totals); below them, each training and sampling request is a clickable slice on its lane.

_An example exported trace_

New to Perfetto?

A general-purpose trace viewer; see the official [Perfetto UI guide](https://perfetto.dev/docs/visualization/perfetto-ui) for the full tour. Essentials:

- **Zoom/pan**: `W`/`S` zoom, `A`/`D` pan (or `Shift`+drag, `Ctrl`+wheel).
- **Click a slice** for its details; `,` and `.` step between slices.
- **Drag a time range** for aggregate stats over it.
- **Find a track**: `Ctrl`+`P` (`Cmd`+`Shift`+`P` on Mac).

_Aggregate stats for a selected time range_

## Metric Details

### Concurrency

- **In-flight sample requests** (`requests`): sampling requests running at each moment (the console's "Active sample requests"). Split by weights version: each colored band is a different generation (e.g. a newer checkpoint), stacking to the total, so you see the checkpoint mix in flight.
- **Training utilization** (`%`): the percentage of wall-clock time the trainer was actively running forward/backward steps in each interval. Intervals are 1 minute long in Perfetto and can be longer in the console. It will not always be 100%; treat any value above 0% as a sign that training work was happening in that interval.
- **Rate-limit events/sec** (`events/s`): how often the session hit a rate or capacity limit, per second.

_The Concurrency counter tracks_

### Throughput

- **Training tokens/sec** and **Generation tokens/sec** (`tok/s`): total training and generated tokens per second. Each request's tokens are spread evenly across the time that request was active (from when it started running to when it finished), instead of being credited all at completion, so a long request reads as a steady rate over its lifetime rather than a spike at the end. **Note that because of this, these numbers are an estimate.** Intervals are 2.5 minutes long in Perfetto and can be longer in the console.
- **Prefill (prompt) tokens/sec** (`tok/s`): total prompt tokens per second, credited at the moment each request starts, because prefill is a burst of work at the very beginning of a request. **Note that because of this, these numbers are an estimate.**

_The Throughput counter tracks_

### Per Request

- **Avg prompt tokens/request** and **Avg generation tokens/request** (`tokens`): cumulative average size across all sample requests completed so far, so it settles toward the session-wide average.
- **Max prompt tokens** and **Max generation tokens** (`tokens`): largest single prompt and generation in each 1-second interval (outliers).

_The Per Request counter tracks_

### Latency

- **Sample request latency** p50/p90/p99 (`s`): sampling request time at each percentile, per 1-second interval.
- **Train request latency** p50/p90/p99 (`s`): same for training requests.

_The Latency counter tracks_

### Totals

- **Total completed samples** (`samples`): running total of finished sampling requests.
- **Total generation / prompt / train tokens** (`tokens`): cumulative tokens generated, prompted, and trained on.

_The Totals counter tracks_

### Per-request spans

Below the counters, each training and sampling request is its own slice on a lane, plus p90/p95/p99 rows that isolate the longest-running samples by span length so the slowest requests stand out. Click one for its timing and to see how requests overlapped; sampling slices are colored by weights version (matching Active sample requests).

_Per-request sample spans_

## Example insights

- **Async RL is working.** Sample request spans overlap heavily with little gap between them, so new requests start before earlier ones finish. That overlap is the goal: the sampler stays busy instead of idling between steps.
- **Higher concurrency lifts throughput.** In the second half of this trace, **In-flight sample requests** and **Generation tokens/sec** track each other: concurrency rising from about 80 to 110 requests lifts throughput from roughly 1,500 to 1,900 tokens/sec.
- **High concurrency with low throughput points to a backend issue.** When **In-flight sample requests** spikes early and stays high while **Generation tokens/sec** and **Prefill (prompt) tokens/sec** stay low, requests are piling up faster than they get served. That points to a Tinker backend bottleneck rather than your code, so [email us](mailto:tinker@thinkingmachines.ai?subject=Tinker%20Session%20Issue&body=Session%20ID%3A%20) with your session ID.

## Tips

- **Re-export to refresh.** Downloading again picks up new activity; if nothing changed, you get the existing trace back instantly.
- **Panels for a glance, trace for a deep dive.** Panels for a quick read; the trace to inspect individual requests or correlate a dip with what was running.
