-
Notifications
You must be signed in to change notification settings - Fork 48
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
perf(cuda/quant): Turing (sm_75) needs its own tensor-core arm
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:lowLow priorityLow prioritystatus:backlogIn the backlog, not yet readyIn the backlog, not yet readytype:performancePerformance improvementsPerformance improvementsStatus: Open.#1654 In lablup/mlxcel;chore(release): refresh Linux aarch64 CUDA artifact for specialized serving
platform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:highHigh priorityHigh prioritystatus:backlogIn the backlog, not yet readyIn the backlog, not yet readytype:choreMaintenance tasks (build, CI, etc.)Maintenance tasks (build, CI, etc.)Status: Open.#1653 In lablup/mlxcel;fix(test): Inkling-VL mixed-prefill scatter test fails intermittently under the full suite and passes in isolation
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1622 In lablup/mlxcel;fix(youtu_vl): multi-window images still described wrong after the patch-order fix
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodeltype:vlmVision-language modelVision-language modelpriority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1618 In lablup/mlxcel;chore: bench_decode.sh all measures duplicate checkpoints twice
area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)priority:lowLow priorityLow prioritystatus:doneCompletedCompletedtype:choreMaintenance tasks (build, CI, etc.)Maintenance tasks (build, CI, etc.)Status: Open.#1615 In lablup/mlxcel;perf(bench): Confirm or clear three small-model prefill drops at 0.6.0
area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)priority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1614 In lablup/mlxcel;fix(youtu_vl): processor ignores preprocessor max_num_patches
area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatamodeltype:vlmVision-language modelVision-language modelpriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1611 In lablup/mlxcel;[nightly-verify] main is red
priority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1599 In lablup/mlxcel;fix(lora): fail the load when staged runtime adapter terms are never claimed
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1577 In lablup/mlxcel;fix(test): mlxcel-core CUDA test binary crashes at the default thread count
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersplatform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:mediumMedium priorityMedium prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1566 In lablup/mlxcel;perf: first-token latency is dominated by lazy weight materialization
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#1564 In lablup/mlxcel;fix(inference): fused_sample_probs differs by 1 ULP at temperature 1.0 on sm_70
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)platform:linuxLinux (CUDA / packaging) specificLinux (CUDA / packaging) specificpriority:mediumMedium priorityMedium prioritystatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#1563 In lablup/mlxcel;