Model labs

Test training and inference configurations at production scale before you run them.

Use Scala when selecting capacity, preparing a training run, or tuning an inference deployment. Simulate the workload at target scale and compare changes to the model, serving configuration, parallelism strategy, placement, and network.

Use Scala for capacity, training, and inference decisions.

Each analysis starts with a calibrated workload trace and the network being evaluated. The result measures the performance and cost of the configurations you choose to compare.

Choose capacity before committing

Simulate a training or inference workload across provider fabrics and cluster shapes. Compare performance and cost at the capacity you are considering.

Prepare each training run

Test node count, placement, sharding, parallelism strategy, and collective configuration before scheduling the production cluster.

Test inference performance and tuning

Compare serving and network configurations by throughput, time-to-first-token, tail latency, and the amount of capacity required.

Run the analysis as models and configurations change.

The first analysis establishes the calibrated workload. Scala’s API can then test new cluster shapes, training configurations, and inference deployments.

Calibrate against a measured run

Capture the workload at an accessible scale and compare the simulation with a result measured on hardware. Use step time for training or the performance metric being tested for inference.

The result records the measured run, calibration tolerance, target topology, and node count used for the projection.

Automate configuration sweeps

Use Scala’s API to test node counts, provider fabrics, parallelism strategies, placements, collective configurations, and inference-serving settings.

Store the selected configuration with the training or deployment plan. Compare the observed production result with the projection after the workload runs.

Review a training or inference workload.

A twenty-minute call covers the capacity, training run, or inference deployment being evaluated; the available calibration run; and the performance or cost metric used to compare configurations. No trace is required for the initial call.