How it works
Simulate how AI workloads perform on the networks you operate, use, or design.
Scala runs a calibrated workload trace against a model of the hardware, topology, and transport. Change the workload, network configuration, or node count to compare performance before buying capacity, deploying a configuration, or finalizing a design.
Use Scala’s API to compare configurations by cost, duration, throughput, time-to-first-token, or another metric you specify.
Capture or select a workload trace.
Capture the workload being evaluated, whether it is yours or a customer’s. Use Scala’s library when a specific workload is not available.
Capture a workload
Capture the communication behavior generated by the model code and parallelism strategy. Calibration compares the trace with a measured step time from a run of that workload.
A captured trace is required when the result needs to describe your workload.
Use Scala’s trace library
Select a captured and calibrated workload when you cannot collect a trace yet or when you need to compare configurations across several workload types.
Replace the library trace with your own when the analysis needs to support a workload-specific decision.
Calibrate the trace and set the target scale.
Compare the trace with the measured run, tune it to the agreed tolerance, and set the node count you need to evaluate.
Establish the baseline
Compare simulated and measured step times and tune the trace until the result is within the specified tolerance. That calibrated run becomes the baseline for the configurations that follow.
Scala’s team supports initial setup and calibration. Repeated simulations can run through the API after calibration is complete.
Set the scale
Scale the calibrated workload to the node count you need to test, from a single node to a production cluster.
This step does not require access to a GPU cluster at the target size.
Describe the network being evaluated.
Configure the hardware, topology, and transport for a network you operate, use, or design.
Define the topology
Specify the node count, node groups, switching layers, and oversubscription at each layer.
Run the same calibrated workload against multiple topologies to compare fabric designs on the same basis.
Configure the hardware and transport
Describe the accelerators and their links, NIC rates, switch buffering, and the transport running across the fabric.
You can compare hardware alternatives before buying or gaining access to them.
Run simulations and compare the results.
Change the scale or configuration and run the workload again. Rank the alternatives by cost, duration, throughput, time-to-first-token, or another specified metric.
See what limits performance
Observe how the run’s time is spent, where bottlenecks appear, and the scale at which each one begins to affect performance.
Each result states its calibration basis and the uncertainty associated with the simulated scale.
Test alternatives through the API
Sweep node counts, topologies, parallelism strategies, and part choices from a script or an existing automated process.
Rank the results by cost, duration, throughput, time-to-first-token, or a metric you define.
Review the inputs required for a Scala simulation.
A twenty-minute call covers the workload trace, calibration run, and network configuration required for the analysis.