Measure how AI knowledge develops.
Watch internal structure, motion, and domain relationships develop as a model trains. Use the measurements to ask more precise questions about what it has learned.
Spot problems earlier in training. We open AI’s black box so you can watch information organize into knowledge and measure learning as it happens.
We build instruments for research and tools for education, helping people investigate how AI learns and understand what its internal patterns reveal.
Watch internal structure, motion, and domain relationships develop as a model trains. Use the measurements to ask more precise questions about what it has learned.
A neural-network investigation lab. Build, change, predict, and observe—with the explanation next to the experiment.
Different layers.
Different internal footprints.
A score tells you an outcome. An internal view helps you investigate what changed along the way. Our approach is to make those observations visible, keep their meaning clear, and leave room to test the explanation.
Try an interactive example →Loss tells you how a model scored. IRIS lets you watch its internal representations develop while it trains, then compare those changes with validation evidence.
An inspection instrument for research teams · Pre-alpha
Keep loss, accuracy, and held-out evaluation. Add an internal view that helps you locate a change and choose what to inspect next.
Compare effective rank across layers and frames. Inspect how broadly measured variation is distributed.
Follow basis rotation and representation motion over time. Place internal changes beside your outcome curves.
Inspect representation overlap and measured agreement or opposition between probe gradients.
Two invented internal patterns share the same loss curve. First compare the definitions below, then switch patterns and move through training.
Several layers retain higher effective rank: measured variation uses more independent directions within each of those layers.
Effective rank rises in one layer while falling in others. The measured pattern becomes more uneven across layers.
These names describe the two patterns in this illustration. Here, effective rank is k95: the whole-number count of leading eigenvalues needed to capture 95% of measured activation variance. Neither pattern alone establishes better learning.
The same improving loss. It cannot tell you which internal pattern is unfolding.
Invented values for teaching, not a captured run. Higher or lower rank is not a quality grade. Check the pattern against your task, probes, and validation.
Real screenshots from the IRIS bench. Inspect a view, read what it measures, and keep the next question attached to the evidence.
Effective rank describes how many directions carry most measured activation variation. Compare trajectories to see which layers changed.
Next question: how do those changes line up with domain-level evaluation?
New diagnostic frames appear automatically in the Bench web interface as training runs, so your team can examine changes without waiting for the run to finish.
Set a fixed extraction interval in training steps, or use adaptive cadence to capture more often when the model is changing and less often when it is quiet. During a live run, adjust the schedule in the Bench or tune adaptive sensitivity.
Use the visual Bench to compare layers, domains, and training frames. Export chart data as Excel workbooks or JSON, or use the optional Butler to relay scalar measurements to W&B, TensorBoard, MLflow, or Aim.
IRIS offers more than 100 built-in views; availability depends on your model and captured data. For another measure, compute it in your training code and report it as a User Defined Signal, as in the MIT-inspired flow-complexity example. You can also combine signals and IRIS measurements in a User Defined Function in the Bench.
Capture cadence, analysis latency, and available diagnostics determine what you can see and when. Measure capture overhead on your workload. Earlier evidence is an opportunity to review; it is not a promise of time or compute savings.
Download, install, add credits, and connect the tracker to a familiar training script.
Sign in or create an account, complete the account setup, and choose the download for your computer.
Open download dashboard ↗Follow the Install IRIS instructions in your dashboard. They guide you through installing the package and saving your run key.
Choose a credit pack in the portal. The smallest pack is 1 credit. Check the quoted run price and your available balance before training.
Add credits ↗In the Bench, open AI Review Prompts → Add IRIS to Training. Give the copied prompt and your training script to your coding assistant.
The four essential calls are: create IrisTracker, check should_extract(step) and call extract(model, step=step) inside the training loop, then call flush() afterward. Add report_loss(loss) if you want a loss curve in the Bench. The prompt can add other available measurements while preserving your model and training logic.
Small experiments start at 0.01 credits per run. Through 10 million parameters, the rate is 0.05 credits per million parameters.
$1 = 1 credit
Pricing counts all model parameters, including frozen ones. Run duration does not set the charge. Prices are rounded down to 0.01-credit steps, with a 0.01-credit minimum.
Add credits ↗Purchased credits expire twelve months after purchase. Check the portal for current pricing and account requirements before you buy or begin a run.
No. IRIS complements performance measurements with internal observations. A structural change does not by itself establish better quality, generalization, or a cause.
No. IRIS reports measurements. It does not prescribe optimizer settings, change model weights, or decide whether a run should stop.
Availability depends on the model, captured channels, probes, and configuration. Some diagnostics require additional measurements; expert views require a mixture-of-experts model.
The tracker runs in your environment. The bench reads captured diagnostic artifacts without reloading the model. Licensing requires a network connection at the start of a run. Review diagnostic artifacts before sharing them.
A neural-network investigation lab for students. Make a prediction, change something, and look at what happens inside.
Educational prototype · For students & educators
w1 × x1 + w2 × x2 + b = y
Our first problem is 2 + 4 = 6. We know the inputs, 2 and 4, and we want an output of 6. Move the sliders to find the weights and bias that make it happen. The process of adjusting the weights and biases for all the neurons in the model is called training.
Fine-tune each value in steps of 0.05:
The weights and bias encode a rule for combining numbers, and settings that solve these three sums exactly also work for other number pairs.
More complex models can memorize examples, so testing unfamiliar inputs helps show whether their knowledge extends beyond the problems they have already seen.
Start with 2 + 4. Move a slider until the output reaches 6.
Open a page to explore the lesson. Each page uses your browser’s own scrollbars; use Back and Next to move through all thirteen pages.
DAISY Chapter 1 adapts the supplied teaching lesson for this site.
Start with a prediction. Move a slider. Compare the result with what you expected. A useful lesson leaves students with an observation they can explain—and another question to try.
DAISY is being explored as a teaching edition of IRIS. The prototype uses educational examples and toy calculations; it is not evidence of performance on a trained model.
Explore the recorded grok-showcase-2 run in the IRIS bench. This is a read-only demonstration, captured September 30, 2026.
The showcase is a separate static app with its own files. For the full desktop view or downloadable data, open it in its own tab.