Skip to content

Add mobile_back_litert: LiteRT compiled-model backend - #1171

Open
anhappdev wants to merge 105 commits into
masterfrom
litert
Open

anhappdev wants to merge 105 commits into
masterfrom
litert

Conversation

@anhappdev

@anhappdev anhappdev commented Aug 18, 2026 •

Copy link
Copy Markdown
Collaborator

Adds mobile_back_litert: a backend for Android and iOS that runs every benchmark on the LiteRT 2.2.0 CompiledModel API - the LLM benchmarks on a dedicated LLM pipeline, stable diffusion on the stable-diffusion pipeline, the vision/NLP benchmarks on a single-model pipeline. Android runs on the GPU and CPU delegates; iOS, added by #1175 (merged into this branch), runs on Metal.

How it works

  • llm_pipeline drives a litert::CompiledModel with explicit TensorBuffers for prefill and decode.
  • llm-1b and llm-1b-instruct default to the GPU delegate with a dedicated GPU model export (llama_q8_ekv3072_litert.tflite). GPU compilation failure falls back to CPU automatically.
  • The larger LLM benchmarks run on the CPU delegate.
  • single_model_pipeline runs the vision/NLP benchmarks on litert::CompiledModel. The GPU accelerator is the default (fp32 model exports, automatic CPU fallback); the CPU delegate choice runs the int8 exports on XNNPACK. Inputs and outputs stage through host buffers so the mlperf harness keeps stable pointers.
  • stable_diffusion_pipeline runs stable diffusion (added in Run stable diffusion on the LiteRT backend #1174). CPU only on Android: the shipped exports are dynamic_int8 (text encoder, diffusion) and dynamic_fp16 (decoder), aimed at CPU/XNNPACK, so a GPU choice would need dedicated fp32 exports. iOS runs it on Metal from rewritten exports (see iOS below). The sub-models bind by name - the CompiledModel signature order is alphabetical, not positional.
  • NNAPI is not offered: it is deprecated since Android 15, and the CompiledModel NPU path needs vendor SDKs plus AOT-compiled models (the Google Tensor SDK is a sign-up-gated beta, G5-only). The LiteRT team's interim advice for Pixel is the GPU delegate. NNAPI-EdgeTPU coverage stays with the pixel backend on the Pixel 9 Pro job.
  • The backend is listed before libtflitebackend, so LLM defaults to LiteRT; on Pixel devices the pixel backend still outranks it for the vision benchmarks.
  • The backend declines to load on the Galaxy M32 (SM-M326B), which does not have the memory for the LLM benchmarks.

Backend coexistence

  • fallback_policy is replaced by a claim_policy chain: per benchmark, the app walks the backends in priority order and stops after the first claimant that is not CLAIM_SHARED.
  • Every backend declares CLAIM_SHARED, so users can pick any capable backend per benchmark. Defaults still follow the priority order, so CI runs are unchanged.
  • This adds one declaration to each vendor settings file - the only vendor-file change in this PR.

Build and packaging

  • LiteRT is an opt-in flavor like the vendor backends: WITH_LITERT defaults to 0, and the unified test APK, release APK/AAB, and Android CLI builds request it explicitly.
  • APK filename letters: LiteRT adds l, and the Pixel letter changes from g to p (unified suffix is now qsmplt).
  • The board archives a litert-only test APK (android-apk-litert-<n>).
  • The Sonar build keeps WITH_LITERT=1: this backend is first-party code and stays analyzed, unlike the proprietary vendor SDKs.

iOS (#1175)

  • iPhone and iPad run nine benchmarks - the six vision/NLP ones, stable_diffusion, llm-1b and llm-1b-instruct - through the same CompiledModel code, all defaulting to Metal with CPU selectable. llm-3b/llm-8b are not offered (llm-1b alone is ~2 GB resident against a ~3.4 GB per-process cap), and the Gemma benchmarks are not claimed on iOS.
  • Metal comes from the prebuilt libLiteRtMetalAccelerator.dylib 2.2.0, wrapped in a framework and dlopened from the backend's own directory. Minimum iOS moves from 13.1 to 15.0 for every backend and the Flutter framework (the app already targeted 15.0).
  • Runner.entitlements gains kernel.increased-memory-limit and kernel.extended-virtual-addressing.
  • Stable diffusion on Metal uses exports rewritten by tools/sd_gpu/convert.py (published as *_litert_v2.tflite; same weights, CPU output bit-identical) to get past Metal codegen faults and fit the memory cap. On Metal the GPU-compiled models stay resident, because releasing them does not return their memory.
  • Not iOS-specific: custom_buffer_teardown.patch is regenerated so hardware buffers are actually released, CreateOutputBuffers sizes buffers after resize (the object_detection crash), and llm-* quick.max_duration goes from 40 to 150 in tasks.pbtxt.
  • Measured on Metal (iPhone 16 Pro in CI / 17 Pro on BrowserStack): llm-1b 25.0 / 38.1 tok/s, stable_diffusion ~75 s / ~36 s per image, all vision/NLP results valid. The full table, the memory analysis and the known limits are in Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0  #1175.

Merged from compiled_model

Toolchain changes

  • Bazel 7.7.0.
  • @litert 2.2.0; org_tensorflow pinned to the commit LiteRT expects. Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0  #1175 moved it from 2.1.5: the 2.1.5 Metal accelerator could not compile the diffusion model, and the prebuilt accelerator must match the runtime version. The bump applies to every backend and platform.
  • rules_apple/apple_support pinned to LiteRT's versions; small patches for the pinned TF revision.
  • Android builds at NDK API 30 (matches minSdkVersion). Building at 33 made absl reference backtrace(), which Android 11 devices lack, so the backends failed to dlopen there.

CI

  • Device jobs always test the unified APK; the integration test decides the backend. On the Pixel 10 Pro it pins every benchmark to the LiteRT backend (deviceBackendOverride), so that job covers LiteRT vision and LLM (GPU) plus stable diffusion (CPU) end to end.

  • The Pixel 9 Pro job keeps the defaults: vision on the pixel backend, llm-1b/llm-1b-instruct on LiteRT.

  • Measured on the Pixel 10 Pro job at 1b9e045 (LiteRT 2.1.5, before Add iOS support to mobile_back_litert (CPU + Metal), upgrade to LiteRT 2.2.0  #1175) - all nine benchmarks on liblitertbackend, All tests passed. Accuracy comes from the CI subset in quick mode, so it is only meaningful against TFLite on the same run, not as a real accuracy score.

    benchmark accel throughput accuracy
    image_classification_v2 GPU 14.74 0.830
    object_detection GPU 21.91 0.3162
    image_segmentation_v2 GPU 19.28 0.3921
    natural_language_processing GPU 1.757 1.000
    super_resolution GPU 8.789 0.3400
    image_classification_offline_v2 GPU 28.31 0.900
    llm-1b GPU 3.595 tok/s 1.00
    llm-1b-instruct GPU 3.593 tok/s 0.25
    stable_diffusion CPU 0.00796 perf-only
  • LLM CPU baselines: 4.9-8.0 tok/s (Pixel 9 Pro) and 7.97-8.14 tok/s (Pixel 10 Pro). Expected throughput intervals are still deliberately wide (smoke test, not a perf gate); they can be tightened now that the GPU and LiteRT-vision numbers have landed.

  • The unified device jobs got a 90-minute budget and a longer pre-run wait in the test: with the GPU delegate default the LLM benchmarks carry two ~1.2 GB models, doubling the download and checksum-validation time.

  • Vendor jobs (qti/samsung/mtk) skip LLM until those backends support it.

  • tf_nnapi_no_mmap_sharing.patch: the rebuilt TFLite passed the model file's fd to NNAPI drivers. SELinux rejects that for models on external storage, so NNAPI fell back to CPU. The patch restores master's behavior; Pixel EdgeTPU throughput matches master again.

  • Windows build moved from Google Cloud Build to GitHub Actions (windows-build-test.yml).

  • MSVC needed three fixes: hide host Android NDK env from repo rules; LiteRT's is_msvc = True rewrite for xla/tsl; protobuf's --define=protobuf_allow_msvc=true.

  • The Linux Android job uses a GCS bazel cache (same pattern as the Xcode Cloud build) instead of actions/cache, whose 10 GB quota was churned by buildkit blobs and whose keys never re-save.

  • Workflow downloads are pinned, sha256-verified, and https-enforced.

  • iOS gets a litert row in both the build matrix and the BrowserStack matrix (iPhone 16 Pro). After master's matrix rework (deps: support Xcode 27 and iOS 27 #1181) the row builds on macos-26.

  • iOS builds ride out known flakes: on a remote-cache error the build retries without the cache; pod install backs off on CDN 429s.

  • The legacy Cloud Build trigger can be disabled once the new workflow is trusted.

Review notes

  • Two changes from 530df21 are reverted as local testing shortcuts: the llm-* tasks in tasks.pbtxt are restored, and tokenThroughput no longer returns double.infinity. Please confirm neither was deliberate.

farook-edev and others added 10 commits May 19, 2026 10:39
…CL binary).

Includes updating LiteRT to v2.1.5 and Bazel to v7.7.0, along with compile-time error fixes for Android.
Also fixed LLM benchmark names in binary and potential mask overflow issue
# Conflicts:
#	.bazelversion
#	WORKSPACE
#	flutter/lib/benchmark/benchmark.dart
#	flutter/lib/ui/home/benchmark_config_section.dart
#	flutter/lib/ui/settings/resources_screen.dart
#	flutter/windows/flutter/generated_plugin_registrant.cc
#	flutter/windows/flutter/generated_plugins.cmake
#	mobile_back_qti/cpp/backend_qti/SnpeExecutor.cc
#	mobile_back_qti/cpp/backend_qti/allocator.h
#	mobile_back_qti/cpp/backend_qti/mlperf_helper.h
#	mobile_back_qti/cpp/backend_qti/qti_c.cc
#	mobile_back_qti/cpp/backend_qti/soc_utility.cc
#	mobile_back_qti/cpp/backend_qti/tflite_c.cc
#	mobile_back_tflite/cpp/backend_tflite/neuron/neuron_backend.cc
#	mobile_back_tflite/cpp/backend_tflite/neuron/neuron_builder.h
Combines the latest master (e073a03, per-benchmark backend selection)
with the LiteRT/compiled-model work from compiled_model (d93f3d2).

Conflict resolutions:
- .bazelversion: 7.7.0 (required by LiteRT/TF deps) over master's 6.6.0
- WORKSPACE: LiteRT 2.1.5 archive + separate org_tensorflow source repo.
  Master's xcode26_compat.patch was dropped from the litert archive: it
  targets tensorflow/lite/kernels/elementwise.cc, a path LiteRT 2.1.5
  does not have (it uses tflite/). Needs re-porting for Xcode 26 builds.
- QTI/neuron sources: kept master's refactors (SnpeExecutor rename,
  backend_utils, Genie support) and re-applied compiled_model's
  tensorflow/core/platform/logging.h -> absl/log/log.h swap, extending it
  to files master added that compiled_model never saw.
- mobile_back_qti/.../tflite_c.cc: accepted master's deletion.
- Dart + generated Windows plugin files: kept master's versions
  (newer dart formatter style, current plugin set).

Reverted two dev-only changes that commit 530df21 carried alongside the
dependency update, as they are regressions in a combined branch:
- flutter/assets/tasks.pbtxt: restored the 6 llm-{1b,3b,8b}[-instruct]
  task definitions it deleted.
- flutter/lib/backend/loadgen_info.dart: restored the real
  tokenThroughput computation (was hardcoded to double.infinity).
@anhappdev
anhappdev requested a review from a team as a code owner August 18, 2026 07:23
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

The C++ and bazel format jobs failed on code carried over from
compiled_model, which never ran CI. Master is clean on both.

clang-format: compiled_model's tensorflow/core/platform/logging.h ->
absl/log/log.h swap left the include blocks unsorted in 6 files; one
loop in llm_pipeline.cc also needed rejoining.

buildifier (6.0.1, matching CI): reformatted 5 files and fixed the three
lint warnings that fail the job:
- WORKSPACE loaded and called python_init_rules twice, verbatim
- mobile_back_tflite BUILD loaded litert_gpu_accelerator_prebuilts but
  never used it (also unused on compiled_model, so nothing is lost)
- tensorflow_source_rules.bzl print() is deliberate (it surfaces the
  patch script's output), so it is marked buildifier: disable=print
  rather than removed

Verified with the exact commands CI runs: clang-format reports 0 errors
repo-wide and buildifier exits 0.
All four build jobs failed fetching @eigen_archive:

  Cannot find patch file:
  external/xla/third_party/eigen3/eigen_ios.patch

tf-eigen.patch has two hunks. The second patches
third_party/xla/third_party/eigen3/workspace.bzl to add
patch_file = ["//third_party/eigen3:eigen_ios.patch"]. The first creates
that patch file, but its +++ line said
third_party/eigen3/eigen_ios.patch while its own diff --git header said
third_party/xla/third_party/eigen3/eigen_ios.patch. The +++ line wins,
so the file landed at the TF root instead of inside the XLA tree.

This was harmless on master: TF 2.18.0 still has a root third_party/eigen3,
and that root copy is the one its build evaluates, so the label resolved
inside @org_tensorflow and found the file. TF 6d40c20c, which LiteRT 2.1.5
pins, removed the root third_party/eigen3 entirely -- eigen now lives only
under third_party/xla. So XLA's copy is what gets evaluated, and its label
resolves inside @xla, where the file was never created.

Point the +++ line at the path the hunk's own header already names.
Android builds failed compiling SdExecutor.cc:

  QnnApiHelpers.hpp:65:43: error: unknown type name 'float32_t'

float32_t is a NEON type. The vendor header uses it without including
arm_neon.h, relying on it arriving transitively; TensorFlow's headers
used to provide that and LiteRT's do not. The header is downloaded at
build time so it cannot be patched in-tree -- include arm_neon.h from
SdExecutor.h instead, guarded on NEON being available.
backend_utils.cc failed to compile:

  error: implicit instantiation of undefined template
  'std::basic_istringstream<char>'

Only <iosfwd> was reaching it. TensorFlow's headers used to pull in
<sstream> transitively and LiteRT's do not, so include it directly.
qti_c.cc uses std::stringstream the same way and is in the same target,
so it gets the include too.

The other files the scan flagged (coco_gen.cc, ifeval.cc, mmlu_gen.cc)
compiled fine in this build, so they still get it transitively and are
left alone.
iOS builds failed before compiling anything:

  plisttool, line 39: %interpreter_args%
  SyntaxError: invalid syntax

rules_apple's plisttool is a py_binary whose bootstrap stub is generated
from a template. TensorFlow pins rules_apple 3.5.1, which predates the
rules_python that XLA now brings in, so the newer template's
%interpreter_args% placeholder is never substituted and the stub is not
valid Python.

Declare the same versions LiteRT 2.1.5 pins for this dependency set,
before the TensorFlow workspace macros so they win: rules_apple 3.22.0,
rules_swift 2.9.0, apple_support 1.23.1, bazel_features 1.43.0.
apple_support is deliberately 1.23.1 rather than TensorFlow's 1.24.5 --
per LiteRT, that is the highest version without missing LC_UUID,
DEVELOPER_DIR or SDKROOT on macOS Tahoe, which matters here since CI
runs Xcode 26.3. All four sha256 values verified against the archives.
Android: the neuron backend compiles tflite_c.cc and llm_pipeline.cc,
which now use the LiteRT compiled-model API, but its deps were never
updated to match //mobile_back_tflite/cpp/backend_tflite:tflite_c, so
the build failed with undeclared inclusions for ~28 litert headers.
Mirror that target's litert deps.

iOS: the rules_apple bump got past plisttool, and the build now fails
generating @coremltools//:mlmodel_proto:

  DataStructures.proto: File not found.

The .proto files import each other by bare filename while protoc runs
with -Iexternal/coremltools. TensorFlow declares coremltools without a
fix for this; LiteRT declares the same archive with a patch_cmds that
rewrites the imports to be repo-relative. Declare it ahead of
TensorFlow's with that patch. Both projects' coremltools.BUILD are
byte-identical, so the vendored copy is TensorFlow's file unchanged.
…size

Android: the neuron target's deps are fixed but its copts were still
missing what //mobile_back_tflite/cpp/backend_tflite:tflite_c sets, so
single_model_pipeline.cc failed:

  no matching function for call to 'TfLiteInterpreterOptionsAddDelegate'
  no known conversion from 'TfLiteDelegate *' to 'TfLiteOpaqueDelegate *'

The shared sources pass a TfLiteDelegate, so this target needs the same
-UTFLITE_USE_OPAQUE_DELEGATE and -fpermissive.

iOS: coremltools protos now generate and the build gets as far as
compiling LiteRT itself, where the CoreML delegate fails:

  no member named 'resize' in 'google::protobuf::RepeatedField<float>';
  did you mean 'Resize'?

Patch that call site. The float16 branch below it operates on a
std::string, where lowercase resize is correct, so it is left alone;
util.cc's resize is on a std::vector and is also fine. Patch verified to
apply against the v2.1.5 source.
The neuron target failed with an undeclared inclusion:

  'external/org_tensorflow/tensorflow/lite/c/common.h'

Our own sources are migrated to tflite/c/common.h; this include comes
from @neuron_delegate's neuron_delegate.h. MediaTek's delegate is built
against TensorFlow's TF Lite C API rather than LiteRT's, so the header
still has to come from @org_tensorflow. Put the dep on the target that
exposes that header so it propagates to consumers.

Note this leaves the neuron backend pulling TensorFlow's tflite/c:common
alongside LiteRT's equivalent. Getting MediaTek's delegate onto LiteRT
headers would be the real fix if those collide at link time.
The neuron target now compiles but fails to link:

  ld.lld: error: duplicate symbol: TfLiteIntArrayCreate
  ... ~20 more TfLite* C API symbols

Depending on @org_tensorflow//tensorflow/lite/c:common to satisfy
neuron_delegate.h pulled in a second implementation of the TF Lite C
API alongside LiteRT's, so every shared symbol was defined twice.

Rewrite that one include to tflite/c/common.h via patch_cmds on the
neuron_delegate archive and depend on @litert//tflite/c:common instead,
so there is a single TF Lite C API in the .so. The two headers declare
the same API; only APUWareUtilsApi.h and neuron_delegate.h are consumed
from that archive, and only the latter has the include.
Patching neuron_delegate.h in the archive fixed our link but broke the
archive's own build:

  neuron_delegate.h:28:10: fatal error: 'tflite/c/common.h' file not found
  (from @neuron_delegate//neuron/java/src/main/native:native)

That archive is built against TensorFlow throughout -- its targets depend
on tensorflow/lite/c, core/api, delegates/utils, kernels and java/jni --
so it cannot be moved onto LiteRT by rewriting one include, and its JNI
.so is a separate shared library that is fine staying as it is.

The only thing our code needs from that header is TfLiteDelegate, and
LiteRT declares it identically. So generate a copy of just that header
with the include repointed, and expose that instead. Quoted includes
search the includer's directory before -I paths, so it shadows the vendor
header for this package only, the archive builds unmodified, and the
backend .so still has a single TF Lite C API.
The generated-header approach did not shadow anything: generated files
live under bazel-out, so the includer's-directory rule never finds them
and -Iexternal/neuron_delegate still won. The vendor header was used
unchanged and the build was back to:

  missing dependency declarations for
    'external/neuron_delegate/neuron/neuron_delegate.h'
    'external/org_tensorflow/tensorflow/lite/c/common.h'

Nothing else in this target provides the tensorflow/lite/c/common.h
path, so instead of fighting include order, add a forwarding header at
that path which includes LiteRT's tflite/c/common.h, and put its
directory on the include path via includes = ["tf_compat"]. The vendor
header resolves it unambiguously, the archive stays unmodified, and the
backend still links a single TF Lite C API.
The forwarding header could not win: Bazel passes
-iquote external/org_tensorflow for this target, and -iquote is searched
before anything includes = [...] can add, so the vendor header's
"tensorflow/lite/c/common.h" always resolved to TensorFlow's copy.
Putting the forwarding header where -iquote . would find it would mean
creating tensorflow/ at the repo root, which would shadow that path for
TensorFlow's own sources too.

Every route through that include is blocked: depending on
@org_tensorflow//tensorflow/lite/c:common duplicates the TF Lite C API at
link, supplying the header chain as files is blocked because
tensorflow/lite/core/c exports only to //tensorflow/lite:__subpackages__,
and patching the archive breaks its own TensorFlow-based targets.

So stop routing through it. Only single_model_pipeline.cc used the vendor
header, under MTK_TFLITE_NEURON_BACKEND, and all it needs is
TfLiteDelegate, which LiteRT declares identically. Declare the delegate's
API against LiteRT's header instead. The declarations are byte-identical
to the archive's, verified by diff; the header records that they must
stay in sync if the pinned sha256 is bumped.
@anhappdev anhappdev changed the title Merge compiled_model into master (LiteRT branch) LiteRT migration: compiled-model LLM pipeline and Android GPU acceleration Aug 19, 2026
mobile_back_tflite, mobile_back_pixel, mobile_back_qti, mobile_back_apple
and flutter/cpp return to their master state and build against
@org_tensorflow//tensorflow/lite again. The LiteRT compiled-model LLM
pipeline moves to a new mobile_back_litert backend that claims only the
llm-* benchmarks, only on Android, with fallback_policy FALLBACK_FILL_GAPS
so every other benchmark stays on the TFLite fallback. liblitertbackend is
listed before libtflitebackend in list.in, so LLM routes to LiteRT when
both are present.
@anhappdev anhappdev changed the title LiteRT migration: compiled-model LLM pipeline and Android GPU acceleration Add mobile_back_litert: LiteRT compiled-model LLM backend Aug 19, 2026
The builtin py_binary leaves %interpreter_args% unfilled in its launcher
under this workspace's rules_python, so every pbtxt2header action fails
with a SyntaxError. Loading py_binary from @rules_python fixes it, the
same way the branch did before the mobile_back_litert restructure.
litert_c.cc includes litert_settings_android.h, the checked-in wrapper over
the pbtxt2header output, following the tflite backend's pattern; it was
missing from the new backend.

Both iOS jobs fail compiling org_tensorflow's CoreML delegate: the protobuf
bundled with TF 6d40c20c renames RepeatedField::resize() to Resize(). Apply
the same one-line patch already used for LiteRT's copy of the delegate.
flutter/cpp:utils pulls tensorflow/lite/tools/evaluation:utils and with it
TensorFlow's TF Lite C API implementation, which duplicates the one @litert
provides, failing the liblitertbackend.so link. llm_pipeline.cc's include of
flutter/cpp/utils.h was dead - nothing from it is referenced.
anhappdev added a commit that referenced this pull request Aug 27, 2026
…r inputs

The three macOS bazel jobs (Android 35m08s, iOS tflite 14m53s, iOS apple
14m48s) were on the same actions/cache plus --disk_cache setup the Linux build
just moved off, and they get nothing from it. Two runs of the Android-macOS job
over the same commit range:

    run 98382919294  Cache not found              2715 processes: 281 internal, 2434 local
    run 98417618611  Cache restored (166 MB)      2715 processes: 281 internal, 2434 local

Identical action counts and zero hits from a populated cache, so this is action
key instability rather than a storage problem. It cannot be the Linux cause:
these workflows have no setup-gcloud step, so nothing puts a per-run uuid on
PATH, and both runs already had --incompatible_strict_action_env. The runner
image was also identical (macos-15-arm64/20260727.0256), which rules out an
image rotation.

Switching to GCS alone would therefore reproduce exactly what #1171 hit -- a
correctly configured cache that hits nothing -- so this does both halves.

Storage: both workflows drop actions/cache and point bazel at GCS, Android
under bazel-cache/macos-android and both iOS backends under bazel-cache/macos-ios
so each backend warms the cache for the other. Since these jobs, unlike the
Linux one, do not otherwise need a GCS credential, BAZEL_CACHE_ARG is chosen at
runtime and falls back to a disk cache when the secret is absent, so
pull_request runs from forks keep working.

Evidence: a new bazel-cache-fingerprint.sh records an allowlisted environment
and toolchain fingerprint before the build and a sha256 manifest of every
fetched external repo after it, stores both in GCS keyed by run number, and
diffs against the previous run of the same job. That is the method that found
the Linux cause, where exactly one environment value and two of ~240k files
differed. It is portable to macOS, which has no sha256sum, no nproc and no GNU
find -printf, and it can never fail a build.

Note that adding setup-gcloud puts a per-run random path on PATH here too. That
is only harmless because .bazelrc now sets --incompatible_strict_action_env.
Ports the dormant stable-diffusion pipeline to the LiteRT 2.1.5
CompiledModel API and wires it in, closing the last benchmark gap on
this backend. Targets `litert`, not master.

### How it works
- Three models (text encoder, diffusion, decoder) share one
`litert::Environment`; 43 invocations per query at the shipped
`num_steps: 20`. Buffers are created once and reused across the
denoising loop.
- **Inputs are bound by signature name, never by position.** The
signature keys are ordered alphabetically, which does not match tensor
order: the text encoder reads `(positions, tokens)` and the diffusion
model reads `(context, latent, timestep_embedding)`. Binding
positionally would swap latent with context - both float32, so it would
produce a plausible but wrong image rather than an error. A missing name
is a hard failure that logs the names actually found.
- CPU only. The shipped exports are `dynamic_int8` (encoder, diffusion)
and `dynamic_fp16` (decoder), aimed at CPU/XNNPACK; a GPU choice would
need dedicated fp32 exports. Despite the filenames, nothing is quantized
at the graph boundary - encoder inputs are int32, everything else
float32. The delegate is labelled `CPU`/`cpu` accordingly, rather than
copying the TFLite entry's `NNAPI`/`npu` labels, which are cosmetic
(that code attaches no delegate).
- No new model exports and no `tasks.pbtxt` change: the existing v5_0
TFLite SD models are plain flatbuffers and load as-is. All four
checksums were verified against the downloaded files.

### Test coverage, and its limit
Stable diffusion joins the LLM benchmarks in `quickRun`. In accuracy
mode loadgen ignores `min_query_count` and runs the whole set - here 150
sequential 20-step generations plus 150 CLIP scoring passes and a 1.71
GB ground-truth download, which is hours on a CPU against a 90-minute
job. `quickRun` keeps the time-bounded performance phase.

The consequence is worth stating plainly: **CI exercises the pipeline
end to end but cannot catch a wrong-but-plausible image**, because
nothing scores the output. Name-based binding plus hard failure on a
missing name is what guards correctness instead.

The QTI-only gate in `canRunBenchmark` is widened for LiteRT by
inverting the condition rather than adding `|| litert`, so QTI behaviour
is bit-identical and SD does not start running on the Samsung and
MediaTek jobs, which would add ~2.8 GB of downloads and an uncapped CPU
run to three more device budgets.

### Defects fixed in the dormant sources
Carried over from the TFLite twin, all pre-existing: `exit(-1)` on
inference failure now returns `MLPERF_FAILURE`; the per-query invoker
leak is gone; the unbounded scan for a token terminator is bounded at 77
and the token tail zero-padded, so no stale data from a previous sample
reaches the model; the cumulative-alpha index is bounds-checked; the
embeddings loader validates its reads before allocating. Dead code
dropped: `get_tensor_index_by_name`, two unused `get_timestep_embedding`
implementations, and a `main()` behind `__TEST_BPE__`.

Note these files were byte-identical to `mobile_back_tflite`'s copies,
so they now diverge - the same defects remain in the TFLite backend and
are worth a follow-up there.

### Verification
Confirmed on a Pixel 10 Pro (run 33053477420). Stable diffusion ran end
to end on `liblitertbackend`: the pipeline initialised, all six tensors
resolved by signature name with no binding failure, and both samples
completed the full 20-step denoising loop - 111.31 s mean latency per
image (111.71 s p90), 0.00895 QPS, 2 samples in 229.9 s. `isResultValid`
is false only because the min-query rule (2 vs 128) is unmet under the
time-bounded run, exactly as the LLM benchmarks behave.

Two incidental fixes rode along, both surfaced by this work:
- `checkAccuracy` dereferenced `accuracyRun` with a null check operator,
which is null in `quickRun`. The LLM benchmarks never hit it only
because their expected-accuracy maps are empty and return early. The run
mode's own `doAccuracyRun` flag is now threaded through from the call
site, so the check is skipped only when the mode genuinely has no
accuracy phase - rather than whenever the field happens to be null,
which would silently stop checking benchmarks that should have one.
- The BrowserStack trigger retry went from 2 attempts to 8. The device
jobs have no `concurrency:` group, so two boards at once exhaust the
account's parallel budget and every loser fails outright with
`BROWSERSTACK_ALL_PARALLELS_IN_USE` - no session created, nothing run.
That cost four device-job failures on this branch that looked like code
failures. `retry_on_exit_code: 9` is specifically the trigger failure,
so this only lengthens how long a job waits for a slot; everything else
still fails fast. A concurrency group would address the cause rather
than the symptom, but it changes scheduling for every board and belongs
in its own change.

Build side: the sources type-check against the pinned LiteRT headers and
compile under Code Analysis with `WITH_LITERT=1`; `dart analyze`, `dart
format`, `clang-format` and `buildifier` are clean; the settings block
parses under `protoc`; the four model checksums match the md5s of the
actually-downloaded files.
@anhappdev anhappdev changed the title Add mobile_back_litert: LiteRT compiled-model LLM backend Add mobile_back_litert: LiteRT compiled-model backend Aug 28, 2026
@freedomtan

Copy link
Copy Markdown
Contributor

@farook-edev to add some changes including Gemma 4 ones.

@freedomtan

Copy link
Copy Markdown
Contributor

@farook-edev text encoder, unet/diffusion model (50 iterations), VAE decoder. Let's try to enable the UNET one on GPU accelerator first. It seems no re-conversion is needed. Let's try it.

anhappdev added a commit that referenced this pull request Sep 8, 2026
…d time (#1173)

Makes the GCS bazel cache from #1171 actually hit across CI runs, and
rations the BrowserStack device tests to the account's 2-session limit.

## bazel cache

Two things stopped the cache from ever hitting:

- **Linux** inherited `PATH` into every action key. `setup-gcloud` puts
a per-run `$RUNNER_TEMP/<uuid>` on `PATH`, so every key changed every
run — the cache held ~19k entries and hit 7 of 4497 actions.
- **macOS** had three jobs sharing one `actions/cache` key over one
path. Only one could write it, so the Android job kept restoring the iOS
job's cache and never saved its own.

Fixes:

- `--incompatible_strict_action_env` in `.bazelrc` pins the action
`PATH`. Repo rules still see the real environment.
- Pin loadgen's build-date stamps in `WORKSPACE` (the only
non-reproducible fetched file).
- Linux uses the GCS remote cache. macOS keeps a per-config disk cache
**in front of** GCS (fast local hits, GCS for cold cases).
- Every cache key is now per configuration, so the collision can't
recur. Fork PRs with no GCS secret fall back to disk-only.
- Fix the CocoaPods key (was hashing an untracked `Podfile.lock`,
collapsed to a constant).

Result — all jobs now hit 100% of cacheable actions (zero local):

| job | before | after |
|---|---|---|
| Linux build | ~67 min | **33 min** |
| macOS Android | 35 min | **23 min** |
| iOS apple / tflite | ~15 min | ~15 min (durability, not speed) |

## BrowserStack

The account allows 2 parallel sessions; runs were failing with
`BROWSERSTACK_ALL_PARALLELS_IN_USE`.

- Per-PR `concurrency` group on the four device jobs: a new commit
cancels the previous run's job for the same device. Never workflow-wide
— builds always finish.
- `max-parallel: 1` on those matrices, so one run can't exceed the
budget alone.
- Device-job `timeout-minutes` 60 → 135, so the built-in retry
(`max_attempts: 2`) can actually run instead of being killed after the
first attempt.

---

Once merged, #1171 can drop its `ci: move the Linux bazel cache to GCS`
commit and rebase.
anhappdev and others added 7 commits September 8, 2026 06:35
# Conflicts:
#	.bazelrc
#	.github/workflows/android-build-test-linux.yml
#	WORKSPACE
Bazel runs genrules through BAZEL_SH with a fixed PATH of
C:\msys64\usr\bin;C:\msys64\bin plus the Windows directories, so the
Strawberry Perl on the runner image is invisible to them. The TFLite
acceleration/configuration schema genrules shell out to perl and the
runner installs msys2 with the base package set only, which ships sed
but not perl, so the build died with "perl: command not found".

Every previous green Windows run restored a warm bazel disk cache and
resolved the whole graph from it (0 local actions), so the genrule never
ran. This run was the first with all caches expired, which exposed it.
…terializes them

The Android Linux job copies build products straight out of bazel-bin.
Bazel 7 defaults to --remote_download_outputs=toplevel, so only the
outputs of targets named on the command line are downloaded. Two of the
copied files are not top-level outputs -- they are merely srcs of a
cc_library that is:

  flutter/android/commonlibs/lib_arm64/libc++_shared.so   (qti, samsung)
  external/neuron_delegate/.../libtensorflowlite_neuron_jni.so (mediatek)

so when their actions come back as remote cache hits nothing writes them
to disk and the copy fails with "cp: cannot stat". Request the producing
targets explicitly.

This was latent until adaa361 made the GCS cache actually hit: before
that every action ran locally and produced the files as a side effect.
The QTI step failed first; mediatek and samsung would have followed.

Also switch the mediatek jni path to ${BAZEL_LINKS_PREFIX}bin/ like every
other entry, instead of hardcoding bazel-bin/.
@farook-edev

Copy link
Copy Markdown
Contributor

gemma4 model files available here

@freedomtan

freedomtan commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

@anhappdev

  • please help check the model checksum issue.

  • please help put gemma 4 models to cloud side and modify the URLs.

@farook-edev

  • accuracy numbers of the gemma 4 (dynamic w8, activations stay with fp32) you converted.
    • what does the "dynamic wi8" mean?
    • tinyMMLU is done to certain extent
    • ifeval, still checking
  • [] performance numbers, too
    • slower than llama 1B.
    • let's try to see if we can get some profiling data (comparing with llama 1b).
  • things to try: the QAT model (int8 <-- well-support, int8, fp8, and others (fp4?) <-- ??)
    • litert-torch doesn't support converting pre-quantized checkpoints.
  • let's try to convert QWen 3.5 0.8B: https://huggingface.co/Qwen/Qwen3.5-0.8B
  • make sure that the pipe can handle llama 1B, gemma 4 E2B, and QWen 3.5 0.8B ones.
    • the pipeline works for both llama and gemma 4 now.

@farook-edev

farook-edev commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

I found this on quantization here

"""Quantization recipe module.

Existing recipes include:

  1. Dynamic quantization recipes:
    - dynamic_wi8_afp32
    - dynamic_wi4_afp32
  2. Weight-only quantization recipes:
    - weight_only_wi8_afp32
    - weight_only_wi4_afp32
  3. Static quantization recipes:
    - static_wi8_ai8
    - static_wi8_ai16
  4. LiteRT-LM recipes:
    - gemma4_mixed48
    - gemma4_mixed48_hr
    - gemma4_mixed48_b32
    - gemma4_mixed48_b64

Naming convention decoder for recipes:
  - 'dynamic': dynamic range quantization (weights quantized statically,
    activations dynamically at runtime).
  - 'wi[N]': weight integer N-bit (e.g., wi8: 8-bit weights).
  - 'c' or 'b[M]': Granularity. 'c' for channelwise, 'b[M]' for blockwise with
    block size M (e.g., b32: 32-blockwise).
  - 'hr' (optional): Hadamard rotations are used for the quantization
    parameters estimation. (typically for better quality at lower bits, see
    algorithms/uniform_quantize/hadamard_rotation.py).
  - 'afp32': activations remain float32 in the model (may be dynamically
    quantized at runtime by setting compute_precision=INTEGER,
    explicit_dequantize=False).
"""

@freedomtan

Copy link
Copy Markdown
Contributor

If are willing to download models and put them to your device manually, then you can try https://github.com/mlcommons/mobile_app_open/actions/runs/34861220129/artifacts/10356801955. Otherwise, we can wait for the automatically downloading one.

The llm-gemma-e2b and llm-gemma-e2b-instruct benchmarks pointed at
local:///gemma4/... with empty checksums, so the models had to be pushed
onto the device by hand and nothing verified what landed there.

Upload the four files to gs://mlperf-mobile-public/litert/gemma4/ -- the
same bucket llm-1b already uses for llama_q8_ekv3072_litert.tflite -- and
point every model_file entry at the https URL with its md5.

The basenames are kept as-is because the model_filename, embedder_filename,
per_layer_embedder_filename and tokenizer_filename custom_settings
reference them verbatim; the gemma4/ subdirectory keeps the generic names
(model_quantized.tflite and friends) from colliding in the flat litert/
namespace.

Uploaded with parallel_composite_upload_enabled=False: composite objects
carry no md5 in GCS metadata, which would have made the checksums here
impossible to verify server-side and broken the pattern set by the
existing llama object.
llmGemmaE2b and llmGemmaE2bInstruct are in BenchmarkId.allIds, so the
integration test enumerates and runs them, but neither
benchmarkExpectedAccuracy nor benchmarkExpectedThroughput had an entry.
checkAccuracy asserts the per-benchmark map is non-null before it reaches
the "skip when there is no expected value" branch, so both benchmarks
failed with "missing expected accuracy map" on every device job.

Accuracy reuses the existing empty _llm map: the LLM benchmarks have no
reference accuracy yet, and these jobs run in quickRun mode where
accuracy_run is null anyway.

Throughput gets its own _llmGemma rather than reusing _llm, because _llm's
2..50 tok/s bounds do not hold here -- llm-gemma-e2b measured 1.7 tok/s on
Pixel 10 Pro, and the instruct variant reported 63.6 and 405.6 tok/s for
the same model. Registering the backend with an empty per-device map keeps
the non-null assertion satisfied while leaving the value comparison
skipped, so CI stops failing without pinning bounds to numbers that are
not yet trustworthy.
@anhappdev

Copy link
Copy Markdown
Collaborator Author

Gemma E2B measured on BrowserStack

Both Gemma benchmarks now download from gs://mlperf-mobile-public/litert/gemma4/ and the checksums verify on device, so the models load and run on real hardware. Everything below is LiteRT / GPU, quickRun.

The only run where all four completed is build 1035. Build 1036 (9588d5ca) reproduced Pixel 10 Pro llm-gemma-e2b at 1.67 tok/s, and its other three were cut short — see the idle-timeout note below.

device benchmark tok/s TTFT queries duration score
Pixel 9 Pro llm-gemma-e2b 5.64 2.26 s 6 64 s 0.667
Pixel 9 Pro llm-gemma-e2b-instruct 63.63 2.24 s 2 136 s 0.125
Pixel 10 Pro llm-gemma-e2b 1.67 3.85 s 2 59 s 0.0
Pixel 10 Pro llm-gemma-e2b-instruct 405.64 3.60 s 2 421 s 0.125

These are not performance numbers. Every run reports is_result_valid: false with is_min_query_met: false — quickRun only gets 2–6 queries in — and there is no accuracy phase (accuracy_run is null), so the score column is the performance run's incidental result over a handful of samples, not an accuracy measurement.

Two things worth a look.

The instruct throughput is not measuring generation rate. Same model and device as the non-instruct row, yet 405 tok/s against 1.67. It also rises as the device gets slower: Pixel 9 Pro 136 s → 63.6, Pixel 10 Pro 421 s → 405.6. 421 s at 405 tok/s implies ~170k tokens from 2 queries. Something is off in the token accounting on the IFEVAL path.

The Gemma benchmarks exceed BrowserStack's idle timeout, which is what the two red device jobs are. IDLE_TIMEOUT=900 in .github/workflows/scripts/browserstack-app-automate.sh, and one Gemma benchmark is a single uninterrupted test-framework call:

  • Pixel 9 Pro llm-gemma-e2b: 890 s in build 1035 (passed), 894 s in build 1036 (killed)
  • Pixel 10 Pro llm-gemma-e2b-instruct: 880 s in build 1036 (killed)

Most of that is the 4.7 GB download plus the LiteRT GPU compile (Replacing 1534 out of 1534 node(s) with delegate (LITERT_CL)); only ~90–140 s is actual benchmarking. Build 1035 cleared the ceiling by ten seconds, which is why this looks intermittent rather than broken. Note the keepalive ping in flutter/integration_test/first_test.dart is a debugPrint — it reaches logcat but does not reset BrowserStack's idle timer, so it does not help here.

Raising IDLE_TIMEOUT would unblock the device jobs, but the download and compile cost is the real constraint.

@Mostelk

Mostelk commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

@farook-edev is it possible to check pipeline with the Gemma 4 E4B

@farook-edev

Copy link
Copy Markdown
Contributor

@farook-edev is it possible to check pipeline with the Gemma 4 E4B

CPU benchmark using cmdline on a desktop computer with 16GB+ memory should be doable, I'm not sure it'll work on Android because of ram limitations (E2B model was hitting 10GB peaks and 8GB sustained RAM usage).

P.S. the memory usage is for the entire phone, where the system uses approximately 4-6GB of memory.

@freedomtan

Copy link
Copy Markdown
Contributor

@farook-edev is it possible to check pipeline with the Gemma 4 E4B

CPU benchmark using cmdline on a desktop computer with 16GB+ memory should be doable, I'm not sure it'll work on Android because of ram limitations (E2B model was hitting 10GB peaks and 8GB sustained RAM usage).

P.S. the memory usage is for the entire phone, where the system uses approximately 4-6GB of memory.

@farook-edev Google AI Edge Gallery already allows running Gemma 4 E4B IT on some Android and iOS devices, please check it out.

@freedomtan

Copy link
Copy Markdown
Contributor

@farook-edev:

  1. ifeval bug fixed, will pushed. accuracy 64 for ifeval
  2. qwen could be converted, but tensor naming are different, changes to the current pipeline needed.

anhappdev and others added 2 commits September 29, 2026 15:26
…T 2.2.0 (#1175)

Adds the Apple half of `mobile_back_litert`: iPhone and iPad run all
nine benchmarks through the same LiteRT v2 CompiledModel code as
Android, on Metal — including `stable_diffusion`, which needs the
**LiteRT 2.2.0** upgrade this branch also carries.

Stacked on #1171 — merge that first. Up to date with `litert`, including
#1174. Supersedes #1178.

## What it adds

- `mlperf_backend_matches_hardware` returned false off Android; it now
serves a new `litert_settings_apple.pbtxt`. Nine benchmarks claimed, all
selecting Metal.
- New `apple_xcframework`, minimum iOS 15.0 — LiteRT's
`LITERT_MIN_IOS_VERSION`, and what the prebuilt accelerator declares.
Sibling backends and the Flutter framework move up from 13.1 to match;
the app already targeted 15.0.
- Metal comes from a prebuilt `libLiteRtMetalAccelerator.dylib` pinned
at 2.2.0, wrapped in a framework (`ITMS-90426` rejects a bare Mach-O
under `Frameworks/`; its `MinimumOSVersion` is read back out of the
dylib, after a stale 14.0 drew `ITMS-90208`) and dlopened via a
directory recovered with `dladdr` — iOS has no dlopen search path and
`native_lib_path` is empty there.
- `Runner.entitlements` gains `kernel.increased-memory-limit` and
`kernel.extended-virtual-addressing`, in a revertible commit of their
own.
- A `litert` row in both iOS CI matrices; `llm-*` `quick.max_duration`
40 -> 150.
- Two fixes that are not iOS-specific: `custom_buffer_teardown.patch`
regenerated with context (its offsets had made `destroy_func`
unreachable, so no hardware buffer was released — host teardown 1840 ->
875 MiB), and `CreateOutputBuffers` now sizes buffers after resize,
which is why `object_detection` crashed and `image_classification_v2`
did not.

## Why LiteRT 2.2.0

The blocker is `stable_diffusion` on Metal. On 2.1.5 the backend emits
shader source that does not compile for the diffusion model's int8
weights — `use of undeclared identifier 'q0'` — with
`AllowSrcQuantizedFcConvOps` both on and off, so it is code generation
rather than an option we set. The text encoder compiles there; the
diffusion model does not.

It cannot be patched. The codegen ships only inside the prebuilt
`libLiteRtMetalAccelerator.dylib` — `ml_drift_delegate/` in the source
archive is TFLite-facing glue with no shader templates, so
`patches/*.patch` has no way to reach it. Nor can the accelerator move
on its own: a 2.2.0 dylib will not load into a 2.1.5 runtime, or the
reverse.

So the whole pin moves, and `org_tensorflow` with it — the GPU delegate
needs `absl/status:status_macros` from the Abseil that TensorFlow brings
in (`WORKSPACE:186`). That coupling is what makes this a toolchain
change for **every** backend and platform rather than a version bump.
Patches go seven -> four, plus five for upstream gaps the old pins hid.

2.2.0 renames the bug rather than fixing it, which is the next section.

## Stable diffusion on Metal

`tools/sd_gpu/convert.py` rewrites the published v5_0 exports so the
delegate accepts them (fold run-time shapes to constants, drop
`BROADCAST_TO`, squeeze rank-5 tensors, lower a stale `SUB` version).

A fifth rewrite makes it *fit*. Constant tensor sharing decides whether
weights are materialised or dequantised in-shader; without it the
diffusion model needs **4409 MiB** against ~2885 available. Enabling it
hit `error: use of undeclared identifier 'scale'` — the backend picks a
scalar weight-dequant template for **per-tensor int8** weights without
declaring the arguments it references. Single-op models place the fault
exactly: per-tensor int8 `FULLY_CONNECTED` fails; per-axis FC and every
`CONV_2D` form compile.

With the codegen out of reach, the model is the only side we control:
the converter re-expresses those weights as per-axis, repeating the
scale they already carry. No weight byte moves, CPU output
bit-identical. Diffusion drops to **1913 MiB**. Outputs publish as
`*_litert_v2.tflite`, leaving the originals for the base branch and
#1174.

## Memory

iOS kills a process past a per-process cap (3376 MB on an 8 GB device)
and `EXC_RESOURCE` cannot be caught. Apple-only; Android untouched.

- Sharing is on. Off, `llm-1b` needs 4497 MiB and diffusion 4409; on,
2055 and 1913.
- **Releasing a GPU-compiled model does not return its memory** — the
weights are Metal buffers libmalloc never owned. On device, releasing
two of three models recovered 69 MiB of ~1268, and rebuilding one
stacked a second copy and crossed the limit. `set_phase` is therefore
CPU-only; on Metal the three compile once and stay resident (1973 MiB of
3013, peak 2588), which also took a 20-step image from 45.5 s to 18.8 s
on a macOS host.
- `llm-3b`/`llm-8b` are never offered: `llm-1b` alone costs ~2072 MiB
resident from a 1229 MiB file.

## Results

Nine benchmarks in one process, no crash or relaunch. The 16 Pro is the
CI run at `b2f9f166`; the two 17 Pros are a BrowserStack run of the same
test package at `3a0ec360`. QPS, except `llm-*` which is tok/s.

| benchmark | 16 Pro | 17 Pro | 17 Pro Max | |
| --- | --- | --- | --- | --- |
| image_classification_v2 | 93.7 | 105.9 | 102.7 | valid |
| object_detection | 247.5 | 285.5 | 268.9 | valid |
| image_segmentation_v2 | 145.2 | 194.4 | 186.6 | valid |
| natural_language_processing | 50.7 | 61.9 | 60.4 | valid |
| super_resolution | 13.5 | 16.0 | 16.2 | valid |
| image_classification_offline_v2 | 99.0 | 140.3 | 137.9 | valid |
| stable_diffusion | 0.0133 | 0.0280 | 0.0274 | `isMinQueryMet: false` |
| llm-1b | 25.0 | 38.1 | 37.1 | `isEarlyStoppingMet: false` |
| llm-1b-instruct | 25.0 | 38.5 | 33.5 | `isEarlyStoppingMet: false` |

Every benchmark selects Metal on all three, with no CPU fallback and no
memory kill. `stable_diffusion` puts 2732 of 2742 diffusion ops on the
GPU: ~75 s/image on the 16 Pro and ~36 s on the 17 Pro, against ~98 s on
CPU. Metal `llm-1b` measured 23.5–26.1 tok/s on the 16 Pro, against 9.27
on CPU and 3.46 on the Pixel 10 Pro GPU, TinyMMLU 41% vs 42%. The same
BrowserStack build passed on the iPhone 17, 17e and Air too. iPad mini:
Metal wins all six valid vision/NLP benchmarks, by 1.6x to 3.0x.

## Known limits

- The converter's equivalence check is one seeded input set per model
through the CPU `Interpreter` — it does not exercise the fp16 Metal
path, and is not a proof for all inputs.
- `stable_diffusion` and `llm-*` are not `isResultValid`
(`isMinQueryMet` / `isEarlyStoppingMet` false; `llm-*` would need ~64
queries, ~11 min each). SD already carries an exemption in
`performance_result_validity.dart`; extending it to `llm-*` is a
scoring-policy call left for a reviewer.
- `llm-1b*` on the iPad mini is unmeasured on Metal, and claimed on
every iOS device regardless — treat a red iPad run as expected. A
tested-device allowlist is the intended fix.
- No iOS device job asserts throughput, for any backend: BrowserStack
reports `iPhone17,1` while `expected_throughput.dart` is keyed on
`iPhone16,2`. Adding it would make inert assertions live for `tflite`
and `apple` too — a separate change.
- No CoreML or ANE path: LiteRT v2 does not expose one. CoreML stays
`mobile_back_apple`'s job.
# Conflicts:
#	.github/workflows/ios-build-test-macos.yml
@sonarqubecloud

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
C Security Rating on New Code (required ≥ A)

See analysis details on SonarQube Cloud

Catch issues before they fail your Quality Gate with our IDE extension SonarQube for IDE

@freedomtan

Copy link
Copy Markdown
Contributor

@farook-edev: qwen <- converted model no fully-delegated to GPU.

  • why op could not be delegated?
  • performance numbers? not yet

ifeval: will be pushed soon.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants