Google AI Edge

LiteRT Performance Leaderboard

Latency and memory of LiteRT models measured with benchmark_model: on Android lab devices through the Developer Device Platform, on macOS with the same release run locally, and on iPhone inside an app. One row per model file, platform, device and accelerator.

The board

Grouped by platform, sorted by median latency within each. Click a column header to sort and a row to see every field.
No rows match these filters.
The board data did not load.

How the rows are measured

One job, one model file, one accelerator. A row is one job of a benchmark session with LiteRT's benchmark_model at the release named in the Runtime column; the Binary field of a row names the build it ran. Android rows are jobs of a litert benchmark --ddp session: the Developer Device Platform runs the prebuilt Android arm64 binary against the model file it pushes to each device in the matrix and pulls the results back. macOS rows are jobs of a run_local.py session: the prebuilt macOS arm64 binary of the same release runs on the Mac the script is started on. iPhone rows come from the app under ../ios/, which builds the same tool from LiteRT's source and runs it on a phone attached to a Mac or, as an XCTest, on a Developer Device Platform iPhone.

What benchmark_model does. Every row ran the schedule the binary logs at start: a warm-up of at least 0.5 s, then 50 runs or at least 1 s, whichever takes longer, up to 150 s. GPU rows add --use_gpu=true; no driver sets a thread count. Runs is the count the run completed. Each row reports p95 and the run count from a single session.

Where each number comes from. Median, Avg, p95, Init and First inference are the latency metrics in results.pb; Footprint is its overall memory footprint. Nodes delegated N/M, the partition count and the delegate name come from runtime_info.pb (tflite.profiling.ModelRuntimeDetails, which a run writes with --model_runtime_info_output_file): M is the node count of the model's primary subgraph, N the nodes the accelerator's delegate replaced, so a GPU row with a low N ran most of the graph on the CPU, and the partition count is the runtime's for the whole graph: the delegate's runs and the runs that stay off it. A session without that file (the iPhone sessions from a Mac so far) keeps the delegate name its log reports, takes N/M from the log line Replacing N out of M node(s) with delegate (X) where there is one, and shows nodes n/a otherwise. Init on the GPU includes the delegate initializing the graph.

Reading the board. Rows are grouped by platform and ordered by median latency inside the group; a comparison holds only within one platform, device, accelerator and task, which the filters above select. Rows are made by running the drivers and are committed here as data: board.json is this board, measurements.jsonl the jobs behind it, and the README beside this page says how a row is made and how to add a model.