Silicon & fabric

Show how your part performs in a cluster you do not have.

Run captured training and inference traffic against a proposed design, or configure a current product for the workload, network, and scale a customer plans to run.

Before it exists, and after it ships.

The same workload-based analysis answers a design question early and a customer question later.

Test a new product before tape-out

Run captured training and inference traffic against each proposed design. Change link rates, buffer depth, topology, transport, or congestion control and measure the effect on workload performance.

Compare delivered performance, scaling limits, and tail latency before finalizing the architecture.

Test existing silicon in a customer environment

Configure the current product for the customer’s workload, device count, topology, transport, link rates, and target scale. Run several configurations against the same calibrated workload.

Compare throughput, duration, time-to-first-token, or another performance requirement specified by the customer.

Use traffic from captured model workloads.

Scala can supply traces from its workload library or build a trace from customer model code. Both sources produce the communication pattern used in the silicon or fabric simulation.

Scala’s workload library

Select captured and calibrated traces that represent the training or inference workload types the design needs to support.

Library traces allow design comparisons before a customer workload is available.

Customer model code

Capture the communication pattern generated by the customer’s model and parallelism strategy, then calibrate it against a measured run.

Re-run the simulation when the workload or deployment configuration changes.

Review a new product or customer deployment.

A twenty-minute call covers the silicon or fabric being evaluated, workload source, network configuration, target scale, and performance metric.