Watch AI
learn.
Live

Spot problems earlier in training. We open AI’s black box so you can watch information organize into knowledge and measure learning as it happens.

IRIS BENCH / CONCEPT SPACE RECORDED RUN
additiveexponentialmultiplicative
Measure how AI knowledge develops.120,000 training steps
72 probes · 4 layers · 3 domainsOpen the full Bench ↗

Deepen understanding of how AI develops knowledge, and share what we discover.

We build instruments for research and tools for education, helping people investigate how AI learns and understand what its internal patterns reveal.

Two ways to look deeper

One shared curiosity.
Different starting points.

For research & model development
IRIS™

Measure how AI knowledge develops.

Watch internal structure, motion, and domain relationships develop as a model trains. Use the measurements to ask more precise questions about what it has learned.

For students & educators
DAISY

Make learning something you can see.

A neural-network investigation lab. Build, change, predict, and observe—with the explanation next to the experiment.

Discover DAISY ↗Teaching prototype
Through the IRIS lensRecorded evidence

Different layers.
Different internal footprints.

Real IRIS domain spectra plotted as polar petals across four layers
Domain spectra from a recorded IRIS run. Each petal gives a different view of internal structure.
The SACCADIC approach

See what changes.
Investigate what it means.

A score tells you an outcome. An internal view helps you investigate what changed along the way. Our approach is to make those observations visible, keep their meaning clear, and leave room to test the explanation.

Try an interactive example →

What would you like
to look into?

Talk with SACCADIC ↗

Measure how
AI knowledge
develops.

Loss tells you how a model scored. IRIS lets you watch its internal representations develop while it trains, then compare those changes with validation evidence.

An inspection instrument for research teams · Pre-alpha

Inside the instrumentLayer structure
Actual IRIS effective-rank trajectories across four layers of a recorded run
Actual bench capture · IRIS v0.94-Beta · Recorded run v2D-100k-w0.20-s2. A measurement to investigate alongside validation.

The score is only
part of the picture.

Keep loss, accuracy, and held-out evaluation. Add an internal view that helps you locate a change and choose what to inspect next.

01 / Structure

Where does activity concentrate?

Compare effective rank across layers and frames. Inspect how broadly measured variation is distributed.

02 / Motion

When do patterns reorganize?

Follow basis rotation and representation motion over time. Place internal changes beside your outcome curves.

03 / Relationships

How do domains interact?

Inspect representation overlap and measured agreement or opposition between probe gradients.

One loss curve.
Two internal stories.

Two invented internal patterns share the same loss curve. First compare the definitions below, then switch patterns and move through training.

Distributed structure

Several layers retain higher effective rank: measured variation uses more independent directions within each of those layers.

Concentrated structure

Effective rank rises in one layer while falling in others. The measured pattern becomes more uneven across layers.

These names describe the two patterns in this illustration. Here, effective rank is k95: the whole-number count of leading eigenvalues needed to capture 95% of measured activation variance. Neither pattern alone establishes better learning.

Interactive illustration · Synthetic data

Illustrative loss curve

321005,00010,00015,00020,000Loss, L (unitless)Training step, s (steps)

The same improving loss. It cannot tell you which internal pattern is unfolding.

What an internal view adds

0 / 20,000A slow, 10-second sweep across an illustrative 20,000-step run. Drag to explore; release to continue. Pause to hold a frame.

Invented values for teaching, not a captured run. Higher or lower rank is not a quality grade. Check the pattern against your task, probes, and validation.

See the evidence
in context.

Real screenshots from the IRIS bench. Inspect a view, read what it measures, and keep the next question attached to the evidence.

Inspect during
training.

New diagnostic frames appear automatically in the Bench web interface as training runs, so your team can examine changes without waiting for the run to finish.

Capture

Choose when to measure.

Set a fixed extraction interval in training steps, or use adaptive cadence to capture more often when the model is changing and less often when it is quiet. During a live run, adjust the schedule in the Bench or tune adaptive sensitivity.

Inspect

Explore and export.

Use the visual Bench to compare layers, domains, and training frames. Export chart data as Excel workbooks or JSON, or use the optional Butler to relay scalar measurements to W&B, TensorBoard, MLflow, or Aim.

Investigate

Ask your own questions.

IRIS offers more than 100 built-in views; availability depends on your model and captured data. For another measure, compute it in your training code and report it as a User Defined Signal, as in the MIT-inspired flow-complexity example. You can also combine signals and IRIS measurements in a User Defined Function in the Bench.

Capture cadence, analysis latency, and available diagnostics determine what you can see and when. Measure capture overhead on your workload. Earlier evidence is an opportunity to review; it is not a promise of time or compute savings.

Get started
with IRIS.

Download, install, add credits, and connect the tracker to a familiar training script.

  1. Download IRIS.

    Sign in or create an account, complete the account setup, and choose the download for your computer.

    Open download dashboard ↗
  2. Install on your training machine.

    Follow the Install IRIS instructions in your dashboard. They guide you through installing the package and saving your run key.

  3. Add credits.

    Choose a credit pack in the portal. The smallest pack is 1 credit. Check the quoted run price and your available balance before training.

    Add credits ↗
  4. Let your LLM connect the tracker.

    In the Bench, open AI Review Prompts → Add IRIS to Training. Give the copied prompt and your training script to your coding assistant.

    The four essential calls are: create IrisTracker, check should_extract(step) and call extract(model, step=step) inside the training loop, then call flush() afterward. Add report_loss(loss) if you want a loss curve in the Bench. The prompt can add other available measurements while preserving your model and training logic.

    Explore the Bench example →
Credits for a run

A penny for your thoughts.

Small experiments start at 0.01 credits per run. Through 10 million parameters, the rate is 0.05 credits per million parameters.

IRIS credits per run through ten million parametersThe per-run charge rises with total model parameters. One million parameters costs 0.05 credits; five million costs 0.25 credits; ten million costs 0.50 credits. The minimum is 0.01 credits. Duration does not determine this charge.0.00.10.20.30.40.50246810Credits per runTotal model parameters (millions)0.050.250.50
Per-run pricing · 2 October 2026. Current rates ↗
1M parameters
0.05 credits
5M parameters
0.25 credits
10M parameters
0.50 credits

$1 = 1 credit

Pricing counts all model parameters, including frozen ones. Run duration does not set the charge. Prices are rounded down to 0.01-credit steps, with a 0.01-credit minimum.

Add credits ↗

Purchased credits expire twelve months after purchase. Check the portal for current pricing and account requirements before you buy or begin a run.

A few practical answers.

Does IRIS replace validation?

No. IRIS complements performance measurements with internal observations. A structural change does not by itself establish better quality, generalization, or a cause.

Does IRIS decide how to change training?

No. IRIS reports measurements. It does not prescribe optimizer settings, change model weights, or decide whether a run should stop.

Will every view be available?

Availability depends on the model, captured channels, probes, and configuration. Some diagnostics require additional measurements; expert views require a mixture-of-experts model.

Where does the inspection happen?

The tracker runs in your environment. The bench reads captured diagnostic artifacts without reloading the model. Licensing requires a network connection at the start of a run. Review diagnostic artifacts before sharing them.

Bring one run.
Look one level deeper.

Ask about IRIS ↗
DAISY · The teaching edition of IRIS

A little curiosity.
A closer look at
how AI learns.

A neural-network investigation lab for students. Make a prediction, change something, and look at what happens inside.

Educational prototype · For students & educators

Try a neuron experiment

One neuron. Three sums.

w1 × x1 + w2 × x2 + b = y

Our first problem is 2 + 4 = 6. We know the inputs, 2 and 4, and we want an output of 6. Move the sliders to find the weights and bias that make it happen. The process of adjusting the weights and biases for all the neurons in the model is called training.

The first sum moving through one neuron Weight one times 2 plus weight two times 4 plus bias should equal 6. Dotted arrows link the three adjustable values in that equation to their sliders. The input circles show 2 and 4; the number at the right is the neuron's current calculated output y. w1 ×2+ w2 ×4+ b=6 2x1 4x2 +Add 3 y
0.5
0.5
0

Fine-tune each value in steps of 0.05:

Weight 10.5
Weight 20.5
Bias0

The weights and bias encode a rule for combining numbers, and settings that solve these three sums exactly also work for other number pairs.

More complex models can memorize examples, so testing unfamiliar inputs helps show whether their knowledge extends beyond the problems they have already seen.

Try to get the first one, then see if you can get all of them

0 of 3 correct
2 + 4 = 3
Target: 6Error: 3
Keep adjusting
1 + 3 = 2
Target: 4Error: 2
Keep adjusting
5 + 5 = 5
Target: 10Error: 5
Keep adjusting

Start with 2 + 4. Move a slider until the output reaches 6.

DAISY Chapter 1.
The interactive lesson.

Open a page to explore the lesson. Each page uses your browser’s own scrollbars; use Back and Next to move through all thirteen pages.

Start Chapter 1 →

DAISY Chapter 1 adapts the supplied teaching lesson for this site.

Room to learn

“What do you think
will happen?”

Start with a prediction. Move a slider. Compare the result with what you expected. A useful lesson leaves students with an observation they can explain—and another question to try.

DAISY is being explored as a teaching edition of IRIS. The prototype uses educational examples and toy calculations; it is not evidence of performance on a trained model.

Bring curiosity
into the classroom.

Ask about DAISY ↗
IRIS / Recorded bench display

Watch a run unfold.

Explore the recorded grok-showcase-2 run in the IRIS bench. This is a read-only demonstration, captured September 30, 2026.

The showcase is a separate static app with its own files. For the full desktop view or downloadable data, open it in its own tab.

Inside the IRIS bench

Enlarged IRIS bench screenshot

Actual recorded-run screenshot · An observation to interpret alongside the task and validation.