GPU and enclave model
tasqnetwork.io/benchmark/gpu-model
Effective mode A overhead by job length
Session setup takes about 120 ms in this model, of which 0.44 ms is cryptography. Short jobs should share a session. Past about a minute the overhead is the steady-state figure. Steady-state range from Zhu et al., arXiv:2409.03992.
Source: M1_modeA_overhead.csv, model_gpu.py
When is mode A cheaper than redundancy?
Mode A is cheaper than mode R while attested GPU time costs less than this multiple of commodity GPU time. For r = 3 the break-even is about 3x, so confidentiality does not have to cost more than redundancy.
Floating-point reordering
| Reordering | Bit-identical rows | Median max rel. diff | Top-1 agreement | Top-5 set agreement | Samples |
|---|---|---|---|---|---|
| reversed blocks | 0% | 3.75e-7 | 100% | 100% | 400 |
| interleaved | 0% | 3.75e-7 | 100% | 100% | 400 |
Changing only the order of float32 accumulation in a two-layer network breaks bit identity on every row while the predicted tokens stay the same. This is why mode R compares outputs within a tolerance instead of byte for byte. It runs on a CPU and stands in for, but does not measure, GPU kernel nondeterminism.
Source: D1_nondeterminism_cpu.csv