All articles

69 Edge AI NPU Companies: Chips, SDKs & YOLO (2026)

Xuban

There is no single NPU market. This non-exhaustive guide tracks 69 companies and platform ecosystems: 51 vendors with deployable chips or platforms, nine NPU IP suppliers, and nine clearly labeled watchlist companies. They span several incompatible classes of accelerator and almost as many compiler stacks. Most promise the same path: export a model, quantize it, compile it, copy a proprietary artifact to a board, and call a vendor runtime. The details decide whether that path takes an afternoon or an engineering quarter.

This guide maps the companies that provide credible computer-vision hardware in 2026. It records the relevant chips, SDK and compiler names, deployment artifacts, public access level, and models the vendor actually documents. It also gives special attention to Amlogic, whose newer ADLA stack is materially different from the older A311D NPU ecosystem.

Research scope: Sources were checked on August 15, 2026. Named-model claims come from first-party product pages, documentation, model zoos, repositories or vendor tutorials unless explicitly labeled otherwise. Advertised TOPS are vendor figures, not independently comparable benchmark results. This is a developer guide, not investment advice or a paid ranking.

Quick navigation: company tables | Amlogic deep dive | vendor profiles | compiled artifact table | access and lifecycle | LibreYOLO roadmap

The most important finding

The hardware is fragmented, but the deployment pattern is remarkably consistent:

PyTorch checkpoint
    -> ONNX or TFLite interchange graph
    -> representative calibration data
    -> vendor quantizer and graph compiler
    -> chip-specific binary or model package
    -> vendor runtime on the target
    -> application preprocessing, decode, NMS and rendering

ONNX explicitly allows runtimes, code generators and hardware implementations, but an ONNX file is only the handoff. It does not mean that every ONNX operator will lower to every accelerator. Static input shapes, supported-operator constraints, quantization rules and host fallback determine what really runs on the NPU.

The software stack is therefore part of the chip. A nominally fast NPU with a gated or brittle compiler can be a worse deployment target than a smaller device with a public toolchain, model zoo, simulator and stable runtime.

What "supported" means in this guide

Vendor literature uses the word support very loosely. This article separates four levels of evidence:

Evidence levelWhat it provesWhat it does not prove
Validated modelThe vendor publishes a chip-specific model entry, result or compatibility tableYour modified checkpoint retains the same accuracy
Official exampleThe vendor provides an end-to-end tutorial or demo for a named modelOther sizes, tasks or generations compile
Compiler capabilityThe SDK imports a framework or exposes supported operatorsA particular YOLO graph works end to end
Marketing or gatedThe product is intended for vision and a private SDK existsPublicly reproducible model compatibility

This distinction matters. "Imports ONNX" is not equivalent to "supports YOLO11 segmentation." A model can parse but fall back to a CPU, compile with incorrect output numerics, overflow accelerator memory, or lose accuracy during quantization.

NPU companies and platforms at a glance

The tables intentionally mix SoCs, discrete accelerators, smart sensors and adjacent GPU/FPGA targets. They compete for the same deployment decision even when the vendors use names such as NPU, AIPU, BPU, KPU, DLA, HTP, TPU or MLA.

Embedded SoCs, smart sensors and industrial processors

CompanyRelevant chips or platformsSDK, compiler and artifactPublicly documented computer-vision evidenceAccess
AmlogicA311D2, S928X, S905X5/S905D5, A311Y3, C308L/C302X/C302X2, T968D4, C305X2 and A123XAMLNN Toolkit and libnnsdk.so, compiled .adla; separate legacy Acuity .nb stack for A311DThe model playground documents YOLOv5/6/7/8/10/11, YOLOX, YOLOE, YOLO-World, PP-YOLOE, segmentation, pose and OBB on A311D2, S905X5, A311Y3, C305X2 and A123X; the other chips listed here are compiler targets, not entries in that support matrixPublic GitHub and wheels; compatible BSP/driver still required
RockchipRK3562/3566/3568, RK3576, RK3588, RV1126BRKNN-Toolkit2, RKNN Runtime, .rknnYOLOv5/6/7/8/10/11, YOLOX, YOLO-World, PP-YOLOE, segmentation, pose and OBB in the RKNN Model ZooPublic GitHub; binary wheels
QualcommSnapdragon and Dragonwing platforms, QCS6490, QCS8550 and Dragonwing IQ-9075QAIRT/QNN, SNPE and AI Hub; QNN context binaries or DLCThe AI Hub model collection includes YOLOv3/5/6/7/8/9/10/11/26, YOLOX, YOLO-World, YOLOR, RF-DETR and detection/segmentation/pose modelsDocs public; SDK/account requirements vary
Texas InstrumentsAM62A, AM67A, AM68A, AM69A and TDA4x; preview TDA54-Q1Processor SDK Edge AI and TIDLThe Edge AI Model Zoo publishes optimized detection, segmentation, pose and classification models for current AM6xA/TDA4 paths; this is not yet a target matrix for preview TDA54-Q1Public GitHub plus Processor SDK
NXPi.MX 8M Plus, i.MX 93, i.MX 95 and acquired Kinara Ara-1/Ara240eIQ with TIM-VX, Vela, Neutron Converter and Ara SDKThe eIQ Model Zoo includes YOLOv4-tiny/v8, NanoDet, CenterNet, FastestDet, SSD Lite and YOLACT assets or recipes; its YOLOv5 entry is documentation-onlyMixed public and account-gated components
STMicroelectronicsSTM32N6x7 variants, including STM32N657/647, with Neural-ART acceleratorSTM32Cube AI Studio and ST Edge AI CoreCurrent services repository lists Tiny YOLOv2, YOLOv5u, YOLOv8, YOLO11, YOLO26 and ST-YOLOX, plus pose and segmentationPublic tools and repositories
InfineonPSOC Edge E83/E84 with Cortex-M55 plus Ethos-U55; separate NNLite accelerator on the M33 sideModusToolbox, DEEPCRAFT Model Converter, TFLite Micro and Arm VelaArchitecture material names MobileNetV1/V2 as representative kernels and the E84 AI kit includes a camera, but no first-party PSOC Edge YOLO validation matrix was foundPublic documentation, SDK components and current E84 evaluation kits
RenesasRZ/V2L, V2M, V2MA, V2H and V2NDRP-AI Translator, DRP-AI TVM and RUHMIThe RZ/V2H model list validates YOLOv5/v8/v11/26 sizes, YOLOX, pose and segmentation variantsPublic GitHub plus board SDK
SonyIMX500 intelligent vision sensor and Raspberry Pi AI CameraEdge-MDT, Model Compression Toolkit and IMX500 packer; .rpk packageThe Raspberry Pi IMX500 zoo includes YOLOv8n, YOLO11n, EfficientDet Lite, NanoDet+ and SSDPublic Raspberry Pi flow; Sony tooling terms apply
AmbarellaCV72/CV75, CV5/CV52 and CV3-AD familiesCooper Developer Platform, CVflow compiler and runtime DAGsPublic model garden names YOLOX-S, RTMDet-nano, DeepLabV3+, TopFormer, OWL-ViT and LLaVA OneVisionPrimarily partner-gated
SynapticsAstra SL1600/SL1680; SL2611/13/15/17/19; separate SR100 MCU familySyNAP .synap on SL16xx; Torq/IREE .vmfb on SL261x; separate Ethos-U55-oriented SR SDKOfficial YOLOv8n/v8s SyNAP benchmark and SL261x Torq YOLOv8 example; SR evidence must be evaluated separatelyDeveloper portal; some downloads gated
MediaTekGenio 360/360P/420/520/720 with NP8 and MDLA 5.3; Genio 510/700/1200 with older NP6/MDLANeuroPilot Converter, ncc-tflite, Neuron Runtime and .dla; ONNX Runtime path on selected newer systemsThe official IoT AI Hub publishes model pages and benchmarks for YOLOv5 and YOLOv8, plus classification and face modelsPublic Yocto docs; key NP8 bundles and Android material require direct-customer/NDA access
AllwinnerV853 vision SoC with 1-TOPS Vivante NPUAcuity/Pegasus conversion and viplite runtimeOfficial V853 material names YOLOv2/3/4/4-tiny/5/5s, RetinaNet, MobileNet, ResNet and face/person networksPublic docs; SDK package download may require account
Canaan/KendryteCurrent K230/K230D; legacy K210/K510 KPUnncase compiler and runtime; .kmodelThe K230 nncase guide covers YOLOv5s compilation/simulation/runtime; the official demo catalog and current CanMV changelog document YOLOv8/11/26 tasksPublic compiler/PyPI; chip plug-in is binary
Alif SemiconductorEnsemble E7 with two ML accelerators including Ethos-U55; E8 with one Ethos-U85 and two Ethos-U55 NPUsArm Vela and TFLite Micro-based deploymentAlif publishes a family-level INT8 YOLO-Fastest face-detector benchmark, but no broad target-by-target modern-YOLO matrixPublic docs and evaluation kit resources
HimaxHX6538 WiseEye2 endpoint AI MCU with Cortex-M55 and Ethos-U55TFLite Micro plus Arm Vela, with CMSIS-NN and reference-kernel fallbackThe official WiseEye2 examples include YOLOv8n detection, pose and classification, YOLO11n detection, face mesh and PeopleNetPublic GitHub, documentation and purchasable partner board
Analog DevicesMAX78000 and MAX78002 AI MCUs with low-power CNN acceleratorsai8x-training, ai8x-synthesis/izer and MSDK; generated C and weightsOfficial material covers face identification/detection, RetinaNet, Visual Wake Words, classifiers and action recognition; no first-party YOLO deployment was foundPublic GitHub tools, documentation and evaluation boards
D-RoboticsRDK X3/X5/Ultra and newer S100-series BPU boardsOpenExplorer/Algorithm Toolchain and hbm_runtime; X5 .bin, current rdk_s .hbmThe current rdk_x5 branch documents YOLOv5/v5u/v8/v9/v10/11/12/13/26, YOLOE, YOLO-World, segmentation, pose, classification, OCR and CLIP for RDK X5; X3 and S-series use separate branchesPublic GitHub; some toolchain packages are platform-specific
AXERAAX650, AX637, AX630C, AX620Q and AX615Pulsar2 compiler and AXEngine; .axmodelCurrent official samples cover YOLOv5/6/7/8/9/10/11/13/26, YOLOX and YOLO-World; tasks and targets vary by chipPublic samples and docs; compiler distributed separately
SOPHGOBM1684/1684X/1688/1690 and CV186X/CV18xxTPU-MLIR and SOPHON runtime; .bmodel or .cvimodelOfficial demos cover YOLOv3/4/5/7/8/9/10/11/12/26, YOLOX, PP-YOLOE, YOLO-World, OBB, segmentation, face, SAM and OCROpen-source compiler plus vendor runtime
Huawei AscendAscend 310/310P/310B and Atlas 200I/300I edge productsCANN, ATC compiler and AscendCL; .omThe official Ascend ModelZoo includes YOLOv2/3/4/5, YOLOX, YOLOR, SSD, RetinaNet, Mask R-CNN and segmentation modelsPublic docs; CANN packages and kernels must match target
CambriconMLU370-X4/S4 in the cited public matrix; MLU270 is legacy hardware not covered by that matrixNeuware and MagicMindThe repository's MagicMind 1.7 matrix lists YOLOv3/4/5/7/8, PP-YOLOE, SSD, RetinaFace, RetinaNet, Mask R-CNN and segmentation modelsHistorical model examples public; current SDK/container access gated
SunplusSP7350/C3V, approximately 4.1 to 4.6 advertised TOPSVivante Acuity NPU Docker, SNNF runtime and .nbOfficial documentation provides YOLOv5 and a custom end-to-end YOLOv8 deploymentPublic source/docs and board ecosystem
ESWIN ComputingEIC7700/EIC7700X and dual-die EIC7702/EIC7702X RISC-V SoCsENNP: EsQuant, current EsAAC compiler, simulator and ESSDK runtime; .model; older docs used ennc-compileMilk-V-hosted ENNP documentation includes an end-to-end YOLOv3 tutorial, MobileNetV2 and ResNet examplesMixed first-party and board-hosted docs; regional downloads
NuvotonM55M1 MCU with Ethos-U55-256TFLite Micro, Arm Vela and NuEdgeWise/NuMLOfficial NuEdgeWise examples name YOLOv8-nano, YOLOX-nano, YOLO Fastest, SSD-MobileNet and classification modelsMostly public MCU toolchain
RealtekAmebaPro2 RTL8735B / AMB82-mini camera platformArduino/FreeRTOS SDK, VoE and NeuralNetwork API; .nbPackaged models include YOLOv3-tiny, YOLOv4-tiny, YOLOv7-tiny, SCRFD and MobileFaceNetRuntime public; custom converter access requires contact
TelechipsTCC7500 / TOPST AI board, advertised 8 TOPSTC-NN-Toolkit / Enlight SDK; .enlight intermediate and compiled deployment bundleAn official TOPST project deployed YOLOv8s; its UFLD v1/v2 conversions failed on unsupported layers and only a split/postprocessing workaround was proposed. YOLOv4 and another YOLOv8 flow appear in community support threadsBoard docs public; compiler/operator downloads often permission-gated
T-HeadTH1520, advertised 4-TOPS INT8, on LicheePi 4AHHB compiler and CSI-NN2/SHL; hhb.bm plus generated code/parametersOfficial board-vendor examples run YOLOv5n/s and MobileNetV2 on the NPUPublic examples; aging toolchain/maintenance uncertainty
SigmaStarCurrent SSU9383CM smart-camera SoC and older IPU familiesCurrent MI_IPU runtime; an older SGS_IPU toolchain produced .sim to sgsimg.imgAn older official SDK postprocessor names SSD and YOLOv1/2/3; public compatibility of those models and artifacts with SSU9383CM is unverifiedCurrent silicon, but publicly cited compiler/model evidence is generation-split
HiSiliconHi3516CV610/DV500, Hi3519DV500 and Hi3403V100 smart-vision SoCsATC-to-.om with NNN/SVP-NNN board runtimeThe target-specific HiSpark matrix, notably for Hi3403 and Hi3591P, names YOLOv3 through YOLO11, pose, segmentation, OBB, OCR and depth entries; it is not validation across every SoC in this rowActive and distinct from Ascend; repository models are labeled non-commercial-use only

Dedicated edge accelerators and modules

CompanyRelevant siliconSDK and artifactPublicly documented model evidenceAccess
HailoHailo-8, Hailo-8L, Hailo-10H and Hailo-15 familyDataflow Compiler and HailoRT; .hefYOLOv3/v4/v5/v6/v7/v8/v9/v10/11/12 and YOLO26, plus YOLOX, DAMO-YOLO, SSD, EfficientDet, segmentation, pose and OBB, with generation-specific model tablesModel Zoo public; compiler through Developer Zone
Axelera AIShipping Metis products; announced Europa and Titania portfolio productsPublic Voyager SDK; alpha Pipeline Builder .axm/.axe, classic .axmodel bundlesValidated YOLOv3/5/7/8/9/10/11/26, YOLOX, YOLO-NAS, OBB, pose, segmentation and many non-YOLO modelsSDK public on GitHub; customer support account-gated
DEEPXDX-M1 and DX-M1MDXNN SDK, DX-COM and DX-RT; .dxnnModel Zoo listed 354 entries under DX-COM 2.4.0/DX-RT 3.4.0 when checked on August 15, 2026, including YOLOv3 through YOLO11 and YOLO26, YOLOX, SSD, EfficientDet, NanoDet, segmentation, pose and OBBPublic developer resources and suite
MemryXMX3 accelerator and multi-chip modulesMemryX SDK, Neural Compiler and runtime; .dfpOfficial examples and release notes cover detection, segmentation and pose, including YOLOv10, YOLO11 and YOLO26Public documentation and SDK
KneronKL520, KL530, KL630, KL720 and KL730 have current compiler docs; KL830 appears in some PLUS APIs but not the current compiler target listKneron PLUS and Model Toolchain; .nef on documented compiler targetsOfficial YOLO workflow centers on Tiny-YOLOv3; other materials cover YOLOv5Public docs; tool packages/downloads vary by chip
SiMa.aiMLSoC and production Modalix 50-TOPS platformPalette, ModelSDK, MLA Compiler and ModelExecutorPublic releases validate YOLOv7/v8, YOLOX, pose/segmentation, DETR, Mask R-CNN and EfficientDetDocumentation public; product SDK is commercial
EdgeCortixSAKURA-II, advertised 60-TOPS accelerator in M.2 and PCIe productsMERA compiler and frameworkVendor describes vision-to-generative-AI support, but does not publish a sufficiently precise current YOLO compatibility matrixCommercial trial/order inquiry; partner-led validation
BrainChipShipping AKD1500 co-processor/M.2 module; legacy AKD1000 platforms; Akida 2 IPMetaTF packages: akida-models, quantizeml, cnn2snn and akida runtimeThe Akida 2 model card names an AkidaNet0.5 YOLOv2 detector, CenterNet, AkidaUNet and face recognition; those results are not automatically AKD1500 validationCore Python packages public; hardware mapping remains target-specific
BlaizeP1600, Pathfinder and Xplorer productsPicasso SDK, NetDeploy and AI StudioVision is a target market, but no auditable public named-model compatibility matrix was foundCommercial/gated
Google CoralLegacy Edge TPU USB, PCIe, M.2 and Dev Board products; separate active open-source Coral NPU IPLegacy Edge TPU Compiler/libedgetpu/PyCoral; new RISC-V Coral NPU exposes IP, RTL/simulation and ELF examplesLegacy examples emphasize SSD MobileNet, classification, DeepLab and MoveNet; the new IP has no public YOLO deployment matrix and is not a drop-in Edge TPU product successorLegacy core repos archived; new Coral NPU IP actively developed
Lattice SemiconductorECP5, iCE40 UltraPlus, CrossLink-NX, CertusPro-NX, Avant-E and Avant-X FPGAssensAI Studio, Neural Network Compiler and configurable accelerator IPThe current compiler table lists YOLOv1, YOLOv5, YOLOv8 and YOLO11 in target-specific modes, plus SSD, MobileNetV2-SSD, ResNet and ENetCompiler/manuals public; IP terms vary, and Advanced CNN Accelerator is commercial
MobilintREGULUS 10-TOPS and ARIES MLA100 80-TOPS acceleratorsqb SDK; INT8 compiler and .mxqPublic model zoo lists YOLOv3/5/7/8/9/10/11/12/26 detection plus segmentation, pose and OBB variantsPublic docs/models; SDK and hardware are commercial
RebellionsActive ATOM+ CA22 and ATOM-Max CA25; EoL ATOM CA02/ATOM+ CA12; ATOM-Lite CA21 tool status unclearRBLN SDK compiler/runtime/profiler; .rblnOfficial support releases name YOLOv3, YOLOv5/6/7/8 families and current YOLOv8 tutorialsDocs public; compiler wheel requires portal credentials
FuriosaAIVision-oriented Warboy Gen1; separate RNGD generationFuriosa SDK; INT8 .enf on Warboy, separate .fxb on RNGDWarboy zoo documents SSD, YOLOv5M/L and YOLOv7-w6-pose; RNGD's roadmap marks YOLOv8m support completed in 2024 Q4 but exposes no detailed current public vision matrixIAM/account access; keep the two generations separate

Adjacent platforms developers compare with NPUs

CompanyHardwareDeployment stackWhy it belongs in the comparison
NVIDIAJetson Orin GPU plus DLA; Jetson Thor GPU without DLAJetPack, TensorRT and DeepStream; .engineExtremely mature vision deployment; official DeepStream tools document YOLOv4/v7/v8/v9/11, including a YOLO11 OBB configuration. Many results use the GPU, and unsupported Orin DLA layers can fall back to GPU.
AMDKria/legacy DPU targets; Vitis AI 6.2 GA for Versal AI Edge Gen1 VEK280/VE2802 and Gen2 VEK385; separate Ryzen AI client NPUsLegacy .xmodel; Gen1 target-bound NPU snapshot; Gen2 compiled cache directory or production .rai packageVitis AI publishes vision material, but legacy DPU, Versal Gen1, Versal Gen2 and Ryzen AI are separate compilation and deployment contracts.
MicrochipPolarFire SoC FPGA with configurable CoreVectorBlox accelerator IPPublic VectorBlox SDK 3.1 and Libero; .vnnx, .hex and .ucomp deployment assetsCurrent tutorials cover YOLOv5n, YOLOv8 detection/classification/OBB/pose/segmentation, YOLOv9t and other vision models; SDK 3.1 is currently scoped to the PolarFire SoC Video Kit.
IntelCore Ultra NPU, CPU and integrated/discrete GPUOpenVINO IR or compiled model cacheOpenVINO provides a public cross-XPU path; use the verified-model matrix NPU columns rather than assuming every OpenVINO YOLO example is NPU-validated.
AppleApple Neural Engine in A11-and-newer A-series chips and M-series chipsCore ML and coremltools; .mlpackage/.mlmodelcA highly accessible consumer NPU through Core ML, but Apple abstracts the exact CPU/GPU/Neural Engine partition rather than exposing a YOLO-specific hardware compiler.
GreenWaves TechnologiesGAP9 with NE16 neural engineGAP SDK, NNTool and AutoTiler-generated codeOfficial repositories include MobileNet SSD, face detection, classification and a dedicated YOLOX people detector. The full SDK is limited to qualified customers.
SyntiantMass-production NDP200 and sampling NDP250 Neural Decision ProcessorsSyntiant SDK and deployment packagesUltra-low-power always-on vision/sensor processors: NDP200 is documented below 1 mW, while NDP250 image recognition is documented below 30 mW. No broad public YOLO matrix was found.

Amlogic: two NPU generations, not one

Amlogic deserves a deeper explanation because information on the web often combines incompatible products. The company has an older VeriSilicon-based path and a newer proprietary ADLA path. They do not use the same compiler, artifact or runtime.

GenerationRepresentative chipsNPU and toolchainCompiled artifactPractical status
LegacyA311D and S905D3, commonly seen on Khadas VIM3/VIM3LVeriSilicon Vivante VIPNano-QI, Acuity toolkit, KSNN and older aml_npu_sdk.nb inside an nbg_unify output directoryExisting boards and demos; separate legacy integration
Current ADLA2C308L/C302X, S928X, A311D2, T968D4, S905X5/S905D5 and C302X2Amlogic ADLA2, AMLNN Toolkit and NNSDK2Target-specific .adla; W8A8 or W8A16Integer-focused current stack
Current ADLA3A311Y3, C305X2 and A123XAmlogic ADLA3, AMLNN Toolkit and NNSDK2Target-specific .adla; W4A8, W8A8, W4A16, W8A16 or W16A16Adds native INT4, FP16 and BF16 capabilities

The A311D datasheet identifies the original 5-TOPS INT8 NPU. The older Khadas VIM3 application page and NPU overview document DenseNet CTC, MTCNN, RetinaFace, YOLOFace, YOLOv2, YOLOv3, YOLOv3-tiny, YOLOv4, YOLOv7-tiny and YOLOv8n; the overview additionally documents YOLOv8n-pose and marks FaceNet deprecated. Those examples are real, but they are evidence for the old Acuity .nb ecosystem, not for current ADLA chips. A311D is not A311D2, and C305X is not C305X2.

There is a second hardware trap. The Khadas VIM4 revision guide says the original V12/A311D2 revision-B board has no NPU; the advertised 3.2-TOPS NPU appears on V13A and later boards using the A311D2-N0D revision-C part. A product name alone is not enough to select a test device.

The current AMLNN and ADLA stack

The current AMLNN Toolkit covers model conversion, quantization, compilation, inference and profiling. Its pinned May 2026 user guide documents importers for ONNX, floating-point or quantized TFLite, TorchScript .pt, and PyTorch 2 ExportedProgram/PT2. TensorFlow, Paddle and Keras paths are marked as planned in the detailed guide even though the repository overview uses broader wording. The compiler emits .adla. Python deployment uses AMLNN or amlnn_edge_toolkit_lite; native C/C++ links libnnsdk.so through nnsdk2.h. Host-side development can use ADB with nnserver; the runtime also targets Android, Buildroot, Yocto and selected Debian/Armbian environments.

The toolkit's published platform identifiers currently include:

Target IDPlatform
001C308L / C302X
002S928X
003A311D2
004T968D4
005S905X5 / S905D5
006C302X2
007A311Y3
008C305X2 / A123X in the detailed documents and model matrix

That target field is not cosmetic. Amlogic requires the .adla, nnsdk2.h, libnnsdk.so, BSP and compiler/runtime generation to match. The pinned Quick Start calls for ADLA driver 2.0.2 or newer, NNSDK 3.0.0 or newer, and NNSDK2 1.0.0 or newer; the Android deployment guide makes the library, header, driver and BSP relationship concrete. Treat .adla as an ABI-bound build artifact, not a portable model.

Quantization also depends on the NPU generation. ADLA2 targets 001 through 006 document W8A8 and W8A16 compilation. ADLA3 targets 007 and 008 add W4A8, W4A16, W8A8, W8A16 and W16A16 combinations, with native INT4, FP16 and BF16 support in the operator guide. Amlogic recommends 200 to 500 representative calibration samples. Its random-data option is explicitly for performance testing, not accuracy-preserving quantization.

Amlogic's documented model coverage

The pinned Amlogic model playground is much stronger evidence than a generic claim of ONNX support. Its published support/example matrix names A311D2, S905X5, A311Y3, C305X2 and A123X. Compiler acceptance of the other target IDs does not prove that this entire model list was validated on them.

TaskDocumented models
ClassificationMobileNetV2 and ResNet50-v2; DINO on ADLA3
Object detectionPP-YOLOE, YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLOv10, YOLO11, YOLOE, YOLO-World, YOLOX and QR detection
Face and gestureRetinaFace and Gesture Recognition
SegmentationDeepLabV3, PP-LiteSeg, YOLOv5-seg and YOLOv8-seg
Oriented detectionYOLOv8 OBB
PoseBlazePose detector/landmark and YOLOv8 pose
OCR and speechLPRNet, PaddleOCR variants, Whisper Tiny and SenseVoice on selected newer chips
Newer transformer and multimodal workDETR, MobileSAM, CLIP/MobileCLIP and selected low-bit LLM/VLM examples on newer platforms

The same repository publishes the following W8A8 model-runtime figures. These are vendor measurements of the neural network on the NPU, not complete camera pipelines.

Model and inputS905X5A311D2A311Y3
YOLOv8n, 640 x 640101.72 FPS95.14 FPS191.06 FPS
YOLOv8s, 640 x 64042.33 FPS42.77 FPS83.08 FPS
YOLOv8m, 640 x 64019.67 FPS19.82 FPS35.30 FPS
YOLOv8l, 640 x 64010.53 FPS10.12 FPS18.37 FPS
YOLO11n, 640 x 64041.14 FPS41.48 FPS62.24 FPS

The caveats are unusually important. Amlogic's operator guide places NMS in software on both ADLA2 and ADLA3. The README defines these as native NPU model-runtime results and excludes preprocessing and postprocessing; whether every memory transfer is included is not disclosed. The YOLO accuracy rows use 300-image COCO subsets rather than the complete validation set, while the classification rows use ImageNet val1000. The tables also report a different A311Y3 YOLOv8n latency from the FPS table under conditions the repository does not reconcile. Clocks, power mode, cooling, thermal steady state, exact SDK/BSP version and test duration are not disclosed. Use the figures to understand one vendor stack, not to rank it against another.

Why Amlogic is strategically interesting

Amlogic combines four unusual properties: a large embedded SoC footprint, a newly public compiler stack, a current model matrix that overlaps heavily with LibreYOLO's task space, and relatively little first-class integration in mainstream training libraries. The current repositories are also young: they appeared publicly in late 2025 and early 2026, with substantive manuals dated May through July 2026. That creates a genuine first-mover opportunity and a corresponding API-stability risk.

A clean integration should use LibreYOLO's deterministic ONNX export as compiler input, require representative calibration images for low-precision builds, produce .adla plus a metadata manifest, and load through an NNSDK2 adapter. Legacy A311D should use a visibly different .nb target because presenting .nb and .adla as one backend would create silent compatibility failures. Amlogic's repositories carry Apache-2.0 top-level licenses, but redistribution rights for bundled wheels, shared libraries, nnserver and Android AAR binaries should be confirmed directly before LibreYOLO mirrors them.

The strongest public deployment ecosystems

The following companies currently provide the clearest combination of obtainable hardware, public technical material and named-model evidence. That does not make them universally faster. It makes their compatibility claims easier to evaluate before buying hardware.

Hailo

Hailo is the reference example of a dedicated vision accelerator ecosystem. Models are parsed into a Hailo Archive, optimized and quantized using representative images, then compiled into a Hailo Executable Format (.hef) file. Applications load that HEF through HailoRT.

The Hailo Model Zoo contains recipes, pretrained models, postprocessing configurations and performance data for detection, segmentation, classification, pose and other tasks. Its YOLO coverage includes YOLOv3/v4/v5/v6/v7/v8/v9/v10/11/12 and YOLO26 alongside YOLOX, DAMO-YOLO, SSD and EfficientDet. The version-pinned Hailo-8 object-detection table is a better compatibility source than a generic product page, while the standalone application documents deployable YOLO pipelines.

There is an important version boundary. Hailo-8 and Hailo-8L use the older Model Zoo 2.x, Dataflow Compiler 3.x and HailoRT 4.x line, while Hailo-10 and Hailo-15 use the current 5.x generation. Recipes, parser behavior and HEFs are not interchangeable just because every device carries the Hailo name.

Hailo-15 also has a distinct application layer. It uses the Vision Processor Software Package and Hailo Media Library rather than the Hailo-8/10H TAPPAS-style host path. The public hailo-camera-apps reference repository was archived on May 3, 2026, while hailo-apps-core remains a separate public application framework. A repository archive is not proof of product EOL, but Hailo-15 projects should confirm the current application package in the Developer Zone.

LibreYOLO's Hailo deployment guide deliberately stops at a static ONNX handoff because the proprietary compiler cannot be bundled as a Python dependency. That is the honest boundary between producing a suitable graph and claiming a tested HEF.

Rockchip

Rockchip has one of the largest hobbyist and commercial embedded Linux footprints, particularly through RK3588, RK3576 and RK356x boards. The public RKNN-Toolkit2 converts and quantizes models on an x86 Linux host, and RKNN Runtime or Toolkit-Lite executes the resulting .rknn file on the board.

The RKNN Model Zoo is unusually explicit about chips, model families and precision. Its current table lists FP16 and INT8 paths for YOLOv5/6/7/8/10/11, YOLOX, PP-YOLOE, YOLO-World and YOLOv5/v8 segmentation variants; YOLOv8 OBB and YOLOv8 pose are listed as INT8-only. It also covers OCR, face and other segmentation networks.

LibreYOLO already has a conservative direct RKNN exporter. It validates four exact detection variants on RK3588, can compare the compiler simulator with ONNX Runtime, and rejects unvalidated families rather than equating a successful build with correct predictions. That narrow support claim is a useful model for every future NPU backend.

Axelera AI

Axelera AI, often misspelled "Accelera," builds the Metis AIPU and sells it in M.2, PCIe and multi-chip products. Metis is the publicly orderable family; Europa and Titania are announced portfolio products, so availability should not be flattened into one product status. The Voyager SDK is public on GitHub and handles model selection, compilation, quantization, execution and application pipelines; customer support remains account-gated.

The public Voyager Model Zoo is among the most informative in the market. It lists model variant, input size, precision, accuracy and measured performance for YOLOv3, YOLOv5, YOLOv7, YOLOv8, YOLOv9, YOLOv10, YOLO11, YOLO26 n/s/m/l/x, YOLOX, YOLO-NAS, OBB, pose and segmentation, together with many classification and dense-prediction models. Publishing quantization loss and accuracy next to speed is more useful than publishing a peak TOPS number alone.

Artifact naming changed with the API generation. The alpha Pipeline Builder compilation flow emits an .axm model and can package a portable pipeline as .axe. The classic pipeline documentation describes .axmodel plus model metadata and a manifest. An integration must detect the installed Voyager generation instead of hard-coding one extension.

Qualcomm

Qualcomm offers enormous deployment reach, but "Qualcomm NPU" hides several interfaces. Developers encounter the Hexagon HTP through QAIRT/QNN, the older SNPE flow, Qualcomm AI Hub compilation services, LiteRT or ONNX Runtime's QNN execution provider depending on the device and product class.

The official Qualcomm AI Hub Models repository is strong model-level evidence. It publishes optimized packages for YOLOv3, YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLOv9, YOLOv10, YOLO11, YOLO26 detection/segmentation/pose, YOLOX, YOLO-World, YOLOR, RF-DETR and many classification, depth and pose models. AI Hub also exposes device-specific profiling, which is important because a model compiled for one HTP generation is not automatically portable to another.

The main integration cost is matrix size: Android versus embedded Linux, QCS ordering parts versus consumer Snapdragon, HTP architecture, QNN/QAIRT version and quantization scheme all matter. Qualcomm's own current pages disagree on Dragonwing IQ-9075 status: the product page labels it Active and its EVK is evaluable, while the Product Longevity Program still labels IQ-9075 Sampling and lists longevity through 2038. The ordering/SKU prefix is QCS9075; Qualcomm advertises 50- and 100-dense-INT8-TOPS configurations plus Ubuntu/Yocto support. Production status should be confirmed for the exact ordering SKU, and a LibreYOLO integration should record that SKU rather than producing a generic folder called "Qualcomm."

DEEPX

DEEPX's DX-M1/DX-M1M accelerators use the DXNN SDK. DX-COM compiles and quantizes the model, DX-RT executes it, and the deployed artifact uses the .dxnn format. The company publishes the DX-AllSuite, application examples in DX-APP, and a large online Model Zoo.

The public catalog reported 354 entries under DX-COM 2.4.0 and DX-RT 3.4.0 when checked on August 15, 2026. Its documented coverage includes YOLOv3 through YOLO11 and YOLO26, YOLOX, SSD, EfficientDet, NanoDet, DAMO-YOLO, YOLO segmentation, pose and oriented boxes. DEEPX is a particularly plausible integration target because the compiler/runtime boundary and deployment artifact are clearly named and the public vision coverage is broad.

Industrial SoCs are several ecosystems, not one

Industrial vendors often have longer product lifecycles and better camera, safety and real-time integration than maker-board SoCs. Their AI stacks can be more complicated because one vendor may ship several unrelated NPU architectures at once.

Texas Instruments

TI's current edge-AI processors include AM62A, AM67A, AM68A, AM69A and related TDA4 devices. The software path combines Processor SDK Linux, the TIDL compiler/runtime, Edge AI TIDL Tools and the Edge AI Model Zoo. TIDL can integrate with ONNX Runtime, TensorFlow Lite and other application runtimes while offloading supported subgraphs.

TI publishes much more than a framework-import claim: the zoo contains converted models and performance metadata for object detection, segmentation, pose, classification, depth and other tasks. TI also warns that model-zoo artifacts are development starting points rather than automatically production-ready assets. That is a valuable caveat often missing from vendor comparisons.

TI also lists TDA54-Q1 as Preview, with up to four C7 NPUs and up to 400 TOPS on that part; the broader TDA5 family is advertised up to 1,200 TOPS. Those are vendor product claims for a new generation, not permission to transfer AM6xA/TDA4 TIDL model validation to TDA54-Q1 before TI publishes a target-specific matrix.

NXP

NXP requires at least four backend descriptions:

NXP targetAcceleratorCompiler/runtime path
i.MX 8M Plus2.3-TOPS VeriSilicon Vivante NPUeIQ with TIM-VX/VX delegate
i.MX 93Arm Ethos-U65Quantized TFLite plus Arm Vela
i.MX 95NXP eIQ Neutron NPUNeutron Converter and eIQ runtime
Ara-1 / Ara240Kinara-derived discrete NPU; Ara240 advertises up to 40 eTOPSAra SDK, with integration into the broader eIQ story

The eIQ Model Zoo contains assets or recipes for YOLOv4-tiny, YOLOv8, NanoDet, CenterNet, FastestDet, SSD Lite, YOLACT and other models; its YOLOv5 entry was added as documentation only. An entry validated for one backend is not evidence for all four. NXP's own YOLO export guidance for i.MX platforms illustrates this target-specific conversion problem. NXP says the monolithic eIQ Toolkit stopped receiving updates after version 1.17 in Q3 2025; current workflows use standalone packages such as eIQ Neutron SDK.

NXP completed its acquisition of Kinara in October 2025. The Ara240 fact sheet reaches up to 40 vendor-advertised eTOPS and reports 313 images per second for YOLOv8n. NXP lists both Ara SDK and eIQ Toolkit on the product surface, but that does not establish artifact interchange with the separate i.MX 95 Neutron path. The 16 GB M.2 module is active while the USB option is preproduction. NXP defines the "e" in eTOPS as "equivalent," not "effective"; it is not another vendor's dense-INT8 TOPS metric and should not be compared directly.

STMicroelectronics

STM32N6x7 devices, including STM32N657 and STM32N647, bring a 600-GOPS Neural-ART accelerator into an MCU-class product; the N6x5 general-purpose line does not include that accelerator. STM32Cube AI Studio and ST Edge AI Core analyze and optimize imported models, while STM32 AI Model Zoo Services provides deployment, optimization, benchmarking and example workflows.

The current object-detection overview documents Tiny YOLOv2, YOLOv5u, YOLOv8, YOLO11, YOLO26 and ST-YOLOX alongside classification, pose and segmentation services. This is not the same performance class as a 200-TOPS PCIe card. It is interesting because camera ingest, inference and control can live in a deeply embedded power and memory envelope.

Renesas

Renesas RZ/V processors integrate DRP-AI accelerators. The software has evolved from DRP-AI Translator and DRP-AI TVM toward RUHMI, powered by EdgeCortix MERA technology. The open RZ/V DRP-AI TVM repository and the RZ/V2H validation list name YOLOv5, YOLOv8, YOLO11 and YOLO26 n/s/m variants, YOLOX, pose and segmentation networks alongside classification models.

Renesas is a credible robotics and industrial target, but the exact board, DRP-AI generation, translator version and CPU-side postprocessing need to be recorded for reproducibility.

Smart sensors and camera-first processors

Sony IMX500

Sony's IMX500 is not a normal application processor. It stacks image sensing and AI processing so inference can occur in the sensor, reducing the amount of image data that must leave the camera. The Raspberry Pi AI Camera makes that architecture unusually accessible.

The official Raspberry Pi IMX500 model repository publishes packaged YOLOv8n, YOLO11n, EfficientDet Lite0, NanoDet+ and SSD MobileNetV2 FPN Lite models. The AI Camera documentation explains conversion, packaging and on-sensor postprocessing. The deployment unit is an RPK package rather than a generic ONNX file, and sensor memory plus operator constraints are central design limits. Raspberry Pi's product page states that this camera will remain in production until at least January 2028.

Ambarella

Ambarella's CVflow processors are deeply established in cameras, drones and automotive vision. Current product families include CV72/CV75 for cameras, CV5/CV52 for higher-performance vision systems and CV3-AD automotive devices. The Cooper developer platform and CVflow compilation/runtime flow form the publicly named software stack.

The public developer model garden names YOLOX, RTMDet, DeepLabV3+, TopFormer, OWL-ViT and multimodal models. However, detailed compiler documentation and downloads remain partner-oriented. An Ambarella integration should therefore begin with a vendor or OEM relationship, not an assumption that a public pip package is available.

Synaptics Astra

Synaptics has three relevant targets. SL1600/SL1680 Astra systems use the SyNAP toolkit, which converts a source model plus a YAML description into a model.synap package. Newer SL2611/13/15/17/19 devices use the Torq platform, an IREE/MLIR-based flow that produces .vmfb. The SR100 series is a separate MCU family built around Cortex-M55 and Ethos-U55 with its own SDK. These artifacts and APIs are not interchangeable.

The official YOLO benchmark tutorial is unusually concrete: it exports YOLOv8 to TFLite, performs asymmetric UINT8 calibration in SyNAP, compiles for SL1680, then runs synap_cli and the object-detection application. It proves YOLOv8n/v8s on SL1680 and describes SyNAP support through YOLO11. A separate SL261x object-detection guide runs YOLOv8 from a Torq .vmfb. These are target-specific official examples, not a claim that every YOLO graph is supported on all three families.

MediaTek Genio

MediaTek deserves a full entry because its current IoT AI Hub now exposes the hardware/software matrix, model packages and benchmark tables that older market surveys could not audit. Genio 360/360P/420/520/720 use NeuroPilot 8 with MDLA 5.3; Genio 510/700 use NP6 with MDLA 3.0; Genio 1200 uses NP6 with MDLA 2.0. MediaTek states that the NeuroPilot and MDLA generation, operator set, compiler and runtime are version-bound for the life of each SoC.

The analytical AI path is TFLite-centered:

PyTorch model
    -> NeuroPilot Converter and representative calibration
    -> quantized .tflite
    -> version-matched ncc-tflite compiler
    -> chip-generation-specific .dla
    -> Neuron Runtime on MDLA

Newer Genio products also expose online LiteRT delegation and an ONNX Runtime CPU/NPU path on supported operating systems. Those are different execution modes from an offline .dla. The software architecture guide and resource matrix make the distinction explicit.

The official analytical model table publishes YOLOv5s and YOLOv8s benchmark entries and conversion guides across multiple MDLA generations. It explicitly does not distribute preconverted YOLO artifacts because of AGPL-3.0 restrictions. Its Quant8, 640 x 640 offline measurements include 5.35 ms and 8.04 ms respectively on Genio 720. The YOLOv8s model page publishes input/output tensors, backend-specific times and warnings about MediaTek custom operations. These figures are vendor benchmarks, not end-to-end camera results.

Access is the limiting factor. Public Yocto documentation is detailed, but the current NP8 all-in-one converter/compiler bundle and much Android material are listed as NDA or direct-customer resources. A normal developer account does not grant the same access as a MediaTek Online customer account. LibreYOLO should therefore treat Genio as technically proven but partnership-dependent for a reproducible bring-your-own-model compiler integration.

China-focused edge AI ecosystems

Several of the most capable and accessible low-cost vision platforms are poorly represented in English-language market lists.

D-Robotics and Horizon-derived BPU platforms

D-Robotics maintains the RDK developer-board ecosystem around BPU accelerators used in RDK X3, X5, Ultra and newer S100-series products. The current rdk_x5 branch of its RDK Model Zoo documents YOLOv5/v5u, YOLOv8/9/10/11/12/13/26, YOLOE, YOLO-World, detection, segmentation, pose, classification, OCR and CLIP paths for RDK X5. X3 uses a separate branch, while the current S-series delivery lives in the main zoo's rdk_s branch; the older rdk_model_zoo_s repository contains historical demos. The modern X5 list is therefore not evidence that every model was validated on X3, Ultra or S100.

OpenExplorer/Algorithm Toolchain imports ONNX or Caffe for PTQ and supports PyTorch-oriented QAT flows. Deployment is platform-specific: current RDK X5 uses .bin through hbm_runtime over libdnn, while the current primary rdk_s flow uses .hbm through a same-named hbm_runtime API over libhbucp; X3 retains its older branch and inference interfaces. Older rdk_model_zoo_s material also describes historical .bin and .hbm paths, so those artifacts must remain paired with their matching branch and toolchain. Unsupported operators may fall back to CPU. The public examples are good, but version pairing between RDK OS, board firmware, toolchain and model binary remains part of the compatibility contract.

AXERA

AXERA's AX650/AX630/AX620 families appear in compact AI cameras and boards such as Sipeed's MaixCAM line. Pulsar2 imports ONNX, performs calibration and precision analysis, and emits .axmodel; AXEngine runs the artifact.

The official AXERA samples cover YOLOv5/6/7/8/9/10/11/13/26, YOLOX and YOLO-World, with task and chip support varying by platform. On supported targets, the current YOLO26 examples include detection, pose, segmentation and OBB. The Pulsar2 conversion examples expose the real work: exact tensor names, output cuts, layout transforms, calibration archives and target hardware settings.

SOPHGO

SOPHGO's TPU-MLIR is one of the more attractive compiler stacks for independent integration because the compiler itself is open source. It imports ONNX, PyTorch, TFLite and Caffe, lowers and quantizes graphs, and produces .bmodel for BM targets or .cvimodel for CV18xx hardware.

The TPU-MLIR project supports FP32, BF16, FP16 and INT8 flows depending on target. The SOPHON demo collection names YOLOv3/4/5/7/8/9/10/11/12/26, YOLOX, PP-YOLOE, YOLO-World, OBB/segmentation variants, SSD, CenterNet, RetinaFace, SAM/SAM2 and OCR. As always, a demo on BM1684X does not validate the same artifact on BM1688 or CV18xx.

Huawei Ascend

Huawei's edge inference line includes Ascend 310/310P/310B silicon in Atlas 200I and 300I products. CANN provides the ATC compiler, AscendCL runtime and chip-specific operator kernels; ATC turns ONNX and other source representations into an offline .om model.

The current Atlas 200I DK A2 documentation and CANN samples provide a YOLOv5 .om route for Ascend 310B. The broader, older Ascend ModelZoo contains YOLOv2/3/4/5, YOLOX, YOLOR, SSD, RetinaNet, Mask R-CNN, pose and segmentation models, but it should not be presented as if every entry were revalidated on 310B. CANN, kernels, firmware and SoC must match. An .om for Ascend310P is not a generic Ascend artifact.

Cambricon

Cambricon's MLU accelerators use Neuware and the MagicMind inference engine. The public MagicMind Cloud repository is pinned to MagicMind 1.7 and an MLU370-X4/S4 matrix; it is historical compatibility evidence, not a current-SDK or MLU270 validation matrix. That repository documents PyTorch, ONNX, Caffe and Paddle paths; TensorFlow framework support was removed in 1.7 even though historical TensorFlow examples remain. Cambricon's current 3-series developer portal lists later overall SDK releases, so the public example repository and the current commercial stack must not be presented as one version.

The official MagicMind cloud model repository publishes an unusually direct MLU370-X4/S4 matrix. It marks YOLOv3, Tiny-YOLOv3, YOLOv4, YOLOv5, YOLOv7, YOLOv8, PP-YOLOE, SSD, RetinaFace, RetinaNet, Mask R-CNN, DeepLab and UNet examples, including whether C++ or Python examples exist. The evidence is strong; the main barrier is SDK and container access through Cambricon channels.

Canaan/Kendryte

Kendryte's current K230/K230D and legacy K210/K510 KPU chips are important at the low-cost board and smart-camera end of the market. The nncase compiler imports ONNX or TFLite, uses calibration data for fixed-point conversion, simulates the result on a host, and emits .kmodel.

The official K230 nncase development guide walks through compiling, simulating and running YOLOv5s. The separate AI demo catalog includes YOLOv8n detection, segmentation and pose artifacts, while the current CanMV changelog adds optimized YOLOv8/11/26 classification, detection, segmentation, OBB and pose examples. The datasheet's YOLOv5s result is the actual published benchmark; the compiler and binary KPU plug-in versions must match the board SDK.

Allwinner

Allwinner's V853 combines camera-oriented media hardware with a 1-TOPS Vivante NPU. The official English NPU guide describes Acuity/Pegasus import, quantization, verification and deployment through viplite. Allwinner's own model pages and YOLOv5 guide name YOLOv2/3/4/4-tiny/5/5s, RetinaNet, MobileNet, ResNet and face/person networks.

This is real computer-vision support, but it remains tied to the older Tina Linux/Vivante toolchain and account-dependent SDK downloads. It belongs in the market map with that limitation visible.

Overlooked Korean, Taiwanese and Chinese platforms

English-language lists often jump from Rockchip to NVIDIA and miss a second tier of real, documented silicon. Some of these ecosystems have broader current YOLO matrices than better-known Western startups.

Mobilint

Korean accelerator vendor Mobilint sells the low-power REGULUS and higher-throughput ARIES MLA100. Its qb SDK accepts PyTorch, TensorFlow, TFLite, ONNX and Keras-oriented inputs, quantizes to INT8 and emits a compiled .mxq artifact. The public vision matrix names YOLOv3/v5/v7/v8/v9/v10/11/12/26 detection plus YOLO segmentation, pose and OBB variants. The ARIES page even publishes model-specific YOLO11s and YOLO26m results. That is strong integration evidence, though SDK/hardware access remains commercial.

Sunplus

Sunplus SP7350/C3V uses a Vivante-derived NPU stack but packages it differently from Allwinner or legacy Amlogic. Its public flow combines Acuity conversion, an NPU Docker environment, SNNF runtime and .nb output. The official custom YOLOv8 guide walks through conversion and deployment rather than merely claiming framework support. Related IP does not make its .nb portable to another Vivante-based SoC.

ESWIN Computing

ESWIN's RISC-V family includes EIC7700/EIC7700X and the dual-die EIC7702/EIC7702X, with EIC7700 appearing in shipping boards such as Milk-V Megrez. ENNP contains EsQuant, the current EsAAC compiler, golden-data generation, simulation and the ESSDK runtime, with compiled .model output; older ENNP material used the ennc-compile name. The current EsAAC page documents ONNX input, so other frameworks need export or conversion to that supported IR. The Milk-V-hosted ENNP overview and YOLOv3 tutorial establish a public end-to-end path, but not a broad current model/accuracy matrix.

Nuvoton M55M1

Nuvoton's M55M1 is an MCU-class Cortex-M55 plus Ethos-U55-256 design rather than a Linux application processor. NuEdgeWise/NuML and Arm Vela prepare quantized TFLite Micro models. The public NuEdgeWise repository names YOLOv8-nano, YOLOX-nano, YOLO Fastest v1.1, SSD-MobileNet FPNLite and multiple classifiers. These small-network examples are credible evidence for the device's memory/power class, not evidence for standard 640-pixel YOLO variants.

Rebellions and FuriosaAI

Rebellions' ATOM line uses the RBLN SDK and compiled .rbln artifacts. Its official support release names YOLOv3 tiny/full/SPP, YOLOv5 n through x, YOLOv6 n through l, YOLOv7 variants and YOLOv8 n through x. The current card support matrix marks ATOM CA02 and ATOM+ CA12 EoL, while ATOM+ CA22 and ATOM-Max CA25 are Active. ATOM-Lite CA21 has a product page but is absent from that SDK matrix, so its current toolchain status is unclear. Documentation is public, but the compiler wheel requires Rebellions Portal credentials.

FuriosaAI must be split by generation. Warboy is the older computer-vision accelerator, with INT8 .enf artifacts and an official model zoo containing SSD, YOLOv5M/L and YOLOv7-w6-pose. RNGD uses a different .fxb stack aimed primarily at LLM/VLM workloads. Furiosa's current roadmap records YOLOv8m vision support as completed in 2024 Q4, but its current public supported-model and performance surfaces remain LLM/VLM-centered and do not expose a detailed YOLOv8m accuracy/performance matrix. That is released-generation evidence, not a Warboy compatibility claim.

Realtek, Telechips, T-Head and SigmaStar

These four platforms are real but have narrower, older or gated bring-your-own-model surfaces:

PlatformReproducible public evidenceLimitation to preserve
Realtek AmebaPro2Arduino/FreeRTOS SDK packages YOLOv3/4/7-tiny, SCRFD and MobileFaceNet .nb modelsCustom model conversion requires registering interest/contacting Realtek
Telechips TOPST TCC7500Official 2026 project converted and deployed YOLOv8sUFLD v1/v2 conversion failed on unsupported layers; the article only proposes a split/postprocessing workaround. Full compiler/operator resources commonly require emailed permission
T-Head TH1520Sipeed's application guide runs YOLOv5n/s with HHB and CSI-NN2/SHLPublic stack exists, but maintenance and version clarity lag current leaders
SigmaStar SSU9383CMCurrent docs establish MI_IPU runtime APIs; older SGS_IPU material documents .sim to sgsimg.img and an older postprocessor names SSD and YOLOv1/2/3No public proof that the older compiler artifacts or named-model paths apply to SSU9383CM

HiSilicon smart vision is not Huawei Ascend

HiSilicon's Hi3516CV610/DV500, Hi3519DV500 and Hi3403V100 camera SoCs expose an ATC-to-.om workflow through NNN/SVP-NNN runtimes. The official HiSpark model zoo names YOLOv3/4/5/6/7/8/9/10/11, YOLO11 segmentation/pose, YOLOv8 OBB/World/segmentation, OCR, depth and TinySAM in a target-specific matrix, notably for Hi3403 and Hi3591P. It does not establish that every entry runs on every smart-vision SoC listed above. Some entries are roadmap items, so released examples and planned support must stay separate, and the repository states that its provided ModelZoo models are for non-commercial use only. Shared ATC/.om terminology does not prove artifact or runtime compatibility with the CANN-based Ascend product line.

Emerging and specialized accelerator companies

MemryX

MemryX MX3 modules use a dataflow architecture and can combine multiple chips. The Neural Compiler accepts ONNX, TensorFlow Lite, Keras and TensorFlow models and emits a DFP package consumed by the accelerator runtime. Its runtime APIs support multi-model and streaming pipelines.

MemryX documentation and examples include optimized YOLO preprocessing and postprocessing for detection, segmentation and pose. The SDK release notes explicitly name YOLOv10, YOLO11 and YOLO26 support. The public material is technically useful, although it lacks one simple model/accuracy/device table comparable to Axelera's.

Kneron

Kneron's KL520/KL530/KL630/KL720/KL730 products span small USB and embedded accelerators. Kneron PLUS covers the application runtime, while the current Model Toolchain 0.33.1 documents conversion, quantization, evaluation, simulation and .nef compilation for those targets. KL830 appears in some PLUS runtime APIs, but not in the current compiler target list; a generic .nef compilation claim should not be transferred to KL830 without vendor confirmation.

The official YOLO example demonstrates the process with Tiny-YOLOv3, and other Kneron material covers YOLOv5 training and deployment. The evidence is credible but narrower than the broad modern-YOLO matrices from Hailo, Axelera, Rockchip or DEEPX.

SiMa.ai

SiMa.ai's MLSoC and production Modalix 50-TOPS platform use the Palette software environment. ModelSDK, the MLA compiler and ModelExecutor cover import, quantization, compilation, profiling and execution. The tools support INT8, INT16 and BF16-oriented flows depending on the workload and product.

The company's release notes name validated YOLOv7, YOLOv8 detection/pose/segmentation, YOLOX and YOLOX segmentation, DETR, Mask R-CNN and EfficientDet pipelines. One current limitation should not be hidden: Palette SDK 2.1 notes say QAT is not functional because of a Python/PT2E quantization regression. The documented workaround is to create the annotated ONNX in SDK 2.0 and compile it with 2.1. Access and procurement remain enterprise-oriented.

The deployment boundary is also documented: the ModelSDK compilation flow emits a compiled .tar.gz and can produce an executable ELF, while MPK Tool packages a deployable .mpk. These outputs remain bound to the MLSoC/Modalix target and Palette release.

EdgeCortix

SAKURA-II is a 60-TOPS advertised accelerator for vision and generative AI, compiled through the MERA software framework. EdgeCortix also supplies MERA technology to Renesas' newer RUHMI flow.

The SAKURA-II product material establishes the hardware and compiler, and M.2/PCIe hardware is offered through trial or order inquiry. No sufficiently precise, current public YOLO compatibility matrix was found. That absence is useful purchasing information: model validation should be requested before committing to hardware.

BrainChip

BrainChip's Akida is a neuromorphic event-domain processor rather than a conventional dense-tensor NPU. Its MetaTF environment consists of installable Python packages for model creation, low-bit quantization, conversion, simulation and hardware execution. BrainChip now sells a shipping AKD1500 co-processor in M.2 form, alongside legacy AKD1000 hardware and Akida 2 IP.

The current Akida 2 model card names an AkidaNet0.5 YOLOv2 detector, CenterNet, AkidaUNet and face-recognition networks. Other support material includes FOMO and event-based workloads. Akida supports unusually low-bit weights and activations, but successful deployment may require architecture-aware conversion rather than treating it as another generic INT8 ONNX target. An AKD1000 or Akida 2 model-zoo result should not be attributed to AKD1500 until an explicit map/fit report validates it.

Blaize

Blaize sells P1600-based Pathfinder and Xplorer products, with Picasso SDK, NetDeploy and AI Studio providing the software path. The product pages clearly target vision and edge AI, but a current, public, model-by-model compatibility matrix is not available.

This guide therefore classifies Blaize as commercial and gated rather than filling the gap with community claims. A vendor-supplied compilation and accuracy report for the intended model should be a prerequisite.

TinyML and endpoint vision

Arm Ethos-U and Alif

Arm licenses Ethos-U55, U65 and U85 NPU IP to chip companies. The Vela compiler takes quantized LiteRT/TFLite graphs and rewrites supported subgraphs into Ethos-U custom operations. Unsupported operations may remain on the CPU, which means "the application runs" and "the whole network runs on the NPU" are different claims.

Real implementations include Ethos-U55 in Alif Ensemble E7, Ethos-U65 in NXP i.MX 93 and Ethos-U85 in newer systems. Alif's E7 AI/ML AppKit documents two ML accelerators and demonstrates the E7 TFLite-to-Vela route; the Ensemble E8 combines one Ethos-U85 with two Ethos-U55 NPUs. Alif also publishes a family-level INT8 YOLO-Fastest face-detector benchmark at 192 by 192 on Cortex-M55 plus Ethos-U55, but not a broad target-by-target modern-YOLO matrix. These memory-constrained systems suit small classification and detection networks, not arbitrary 640-pixel detectors merely because they are expressible in TFLite.

Infineon PSOC Edge, Himax WiseEye2 and Analog Devices MAX7800x

Infineon's PSOC Edge E83 and E84 families pair Cortex-M55 with Ethos-U55 and add a separate NNLite block on the low-power M33 side. The E84 architecture manual documents 8-bit weights, 8- or 16-bit activations, and a Vela flow in which supported operators become an NPU command stream while unsupported work remains on the CPU. The DEEPCRAFT Model Converter accepts TFLite, Keras and PyTorch paths, while the architecture material names MobileNetV1/V2 as representative kernels rather than target benchmarks. The E84 AI Kit includes a camera and uses ModusToolbox, but the public material does not yet provide a named YOLO compatibility matrix. This is a real programmable vision platform with maturing model-level evidence, not a validated general-purpose detector target.

Himax's HX6538 WiseEye2 also combines Cortex-M55 and Ethos-U55 for always-on vision. Its public software story is unusually concrete for this power class: the official Grove Vision AI Module V2 examples include YOLOv8n object detection, pose and gender classification, YOLO11n detection, face mesh and PeopleNet. The unified path uses TFLite Micro and Vela, then falls back through CMSIS-NN and reference kernels when necessary. These are deliberately tiny models and resolutions, not evidence that a conventional server-size YOLO graph fits.

Analog Devices takes a different route with the MAX78000/MAX78002 CNN accelerators. The public ai8x-synthesis tool quantizes a trained network and generates target-specific C code, weights and accelerator configuration rather than a portable ONNX runtime package. First-party evidence includes face identification, TinierSSD QR detection in the ADI model zoo, and image classifiers. The devices are relevant to endpoint vision, but no current official modern-YOLO matrix was found.

GreenWaves GAP9

GAP9 combines RISC-V compute with the NE16 neural engine. NNTool imports and quantizes networks, while AutoTiler and the GAP SDK generate deployable code. The public GAP9 neural-network menu contains MobileNet, EfficientNet, ResNet, MobileNet SSD, face and recognition applications. GreenWaves also publishes a YOLOX people-detection project with accuracy and simulation results.

This is unusually transparent TinyML evidence, but the complete SDK is available only to qualified customers and the models are substantially smaller than mainstream 640 by 640 YOLO variants.

Lattice sensAI

Lattice sensAI is an FPGA-based alternative to a fixed NPU. sensAI Studio and the Neural Network Compiler map a supported graph onto configurable accelerator IP for ECP5, iCE40 UltraPlus, CrossLink-NX, CertusPro-NX, Avant-E and Avant-X devices. The output is tied to the selected FPGA, accelerator mode, memory layout and bitstream rather than being a portable runtime model.

The current Neural Network Compiler manual includes an unusually direct topology table. It lists YOLOv1 for ECP5, YOLOv5 and YOLOv8 for optimized/advanced CrossLink-NX and CertusPro-NX modes, and YOLO11 only with Advanced mode IP. It also lists SSD, MobileNetV2-SSD, ResNet, ENet, SqueezeDet and other vision networks, with unsupported cells visible by FPGA family. The Advanced CNN Accelerator targets Avant-E/Avant-X/CertusPro-NX and is commercial IP. This is compiler-table evidence, not a promise that one bitstream supports every listed topology simultaneously.

Microchip VectorBlox

Microchip's public VectorBlox SDK 3.1 preprocesses and compiles fully quantized INT8 TFLite networks for configurable CoreVectorBlox IP in PolarFire SoC FPGAs. tflite_preprocess and vnnx_compile produce .vnnx, .hex and .ucomp deployment assets, and the repository includes a simulator and target-side C driver. Upstream TensorFlow, ONNX and OpenVINO graphs still have to pass through the documented TFLite/INT8 preparation path.

The current tutorial matrix includes YOLOv5n, YOLOv8n detection, classification, OBB, pose and segmentation, YOLOv9t, sparse/compressed YOLO variants, MobileNet, EfficientNet-Lite0, ResNet18, MiDaS, FFNet and QuickSRNet. The C postprocessing guide separately names YOLOv2/3/4/5 and Ultralytics detection, pose and OBB decoders. SDK 3.1 currently supports the PolarFire SoC Video Kit and explicitly not the non-SoC PolarFire Video Kit, so older board demos must not be mistaken for the current target contract.

The SDK is publicly downloadable under a Microchip-specific license that restricts the software and derivatives to Microchip products. Microchip offers free but request-based Libero Silver and CoreVectorBlox licenses. Deployment is still an FPGA system contract: accelerator configuration, bitstream, firmware, SDK-generated network data and memory layout must agree. A successful VectorBlox compile does not turn the model into a binary that can be copied to an arbitrary PolarFire design.

Syntiant

Syntiant's NDP200 targets always-on vision, audio and sensor inference below the power envelope of normal Linux SoCs. The current hardware table lists NDP200 as mass production and NDP250 as sampling. The NDP200 product brief documents CNN, RNN and fully connected network support below 1 mW; the NDP250 announcement instead describes always-on image recognition below 30 mW. "Sub-milliwatt" should not be generalized across both parts.

The public material does not establish broad YOLO compatibility, so the product should be evaluated for small always-on classifiers and detectors through a vendor-supported model assessment rather than a generic ONNX export assumption.

Google Coral: two unrelated generations

The legacy Edge TPU products use the Edge TPU Compiler, libedgetpu and PyCoral with fully INT8 TFLite models. Their main Edge TPU and PyCoral repositories are archived. That is a real maintenance risk, but it is not a formal hardware EOL notice.

Google also maintains a separate, active, open-source Coral NPU: RISC-V accelerator IP for integration into low-power SoCs. The public project currently exposes RTL, hardware simulation, a toolchain and ELF examples. It is not a drop-in USB/M.2 successor to Edge TPU and does not publish a current YOLO deployment matrix.

Adjacent GPU, client-NPU and adaptive-compute traps

NVIDIA, AMD, Intel and Apple often appear in NPU searches, but they do not expose one interchangeable class of accelerator.

NVIDIA Jetson: Jetson Orin combines a CUDA GPU with dedicated DLA cores. TensorRT can build a serialized engine for DLA, but unsupported layers run on the GPU only when GPU fallback is explicitly enabled. The current DeepStream YOLO table covers YOLOv4/v7/v8/v9/11, but most entries are GPU targets; its explicit ONNX DLA entry is YOLOv8s INT8. NVIDIA's DLA restrictions explain the operator limits. Thor is another distinction: NVIDIA's DriveOS migration guide says DLA cores were removed from Thor in favor of GPU scheduling flexibility. "Runs on Jetson" therefore does not identify the accelerator.

AMD Vitis AI: The older Kria/DPU generation compiled .xmodel graphs executed through VART. Current Vitis AI 6.2 is GA for two separate Versal AI Edge paths: Gen1 on VEK280/VE2802 uses ONNX, AMD Quark and a target-bound snapshot/subgraph flow; Gen2 on VEK385 has its own compilation and runtime contract. Gen1 publishes YOLOv5, YOLOv7, YOLOv8 and YOLOX examples. Ryzen AI is a fourth, separate client-NPU stack. Do not transfer artifacts or validation among legacy DPU, Versal Gen1, Versal Gen2 and Ryzen AI merely because all use the Vitis or AMD name.

Intel OpenVINO: OpenVINO can target CPU, GPU and Core Ultra NPU through one API. The Intel NPU plug-in documentation identifies supported NPU generations, while the verified-model matrix exposes device-specific columns. A YOLO notebook proves an application recipe, not NPU validation unless the matrix says so. Supported operations, shapes, precision and the installed NPU driver still determine whether the complete graph runs there. OpenVINO IR (.xml plus .bin) is reusable source-level IR; a compiled-model cache is device and software specific.

Apple Core ML: coremltools converts models into .mlpackage, which Xcode compiles to .mlmodelc. Core ML can schedule work across CPU, GPU and Apple Neural Engine, but it intentionally abstracts the exact partition. That makes Apple an excellent deployment platform and a poor candidate for claims such as "the whole YOLO graph ran on the NPU" unless profiling evidence establishes it.

The NPU IP companies behind the chip companies

Some companies do not sell a board or packaged accelerator at all. They license NPU designs to SoC vendors. They matter because the same underlying architecture can appear under several chip brands, but the final SDK and artifact may still be customized by the licensee.

IP supplierNPU family and softwareWhy it matters
ArmEthos-U55/U65/U85 and VelaAppears in NXP, Alif, Infineon, Himax and other embedded products; quantized TFLite/TOSA-oriented flow
Arm ChinaZhouyi AIPU and Compass SDKPublic model zoo names YOLOv1-tiny/v2/v3/v4/v5, YOLOX and YOLOv8-seg; entries are reference models, not proof for every licensee chip
VeriSiliconVivante VIP family and Acuity SDKRelated Vivante NPU families appear in multiple SoCs, including the separately sourced A311D and i.MX 8M Plus paths; each chip vendor still supplies its own integration
Imagination TechnologiesPowerVR Series3NX AX3146/AX3386/AX3596; historical IMG Series4; Neural Compute SDK/IMG DNNLicensable NNA IP rather than a retail board; Open Access exposes 1/5/10-TOPS Series3NX configurations, while an official NC-SDK program names PP-YOLOE, EfficientNet and HRNet but not every IP configuration
CEVANeuPro-Nano/NeuPro-M and NeuPro StudioLicensable NPU IP with import, quantization, compression, graph compilation, simulation and C/C++ generation
MIPSARC NPX6 and MetaWare MX, formerly Synopsys ARCScalable NPU IP and NN SDK acquired by GlobalFoundries and placed in the MIPS portfolio in June 2026; not itself a retail deployment target
CadenceNeo NPU, Tensilica DSPs and NeuroWeave SDKCompiler/interpreter stack used by SoC licensees, with network and quantization support tied to licensed configurations
QuadricChimera GPNPU IP and SDKProgrammable NPU/DSP-style IP licensed into other companies' chips; the final target is the licensee's silicon
ExpederaOrigin NPU IP and software stackLicensable edge NPU architecture, not a directly purchasable LibreYOLO board

Imagination illustrates why access and lifecycle labels belong beside architecture names. Its Open Access program offers three silicon-proven Series3NX configurations with no upfront IP license fee, but only after company evaluation and with paid support/maintenance plus production royalties. The Neural Compute SDK supplies compilation, optimization, quantization and runtime tooling; an official Baidu PaddlePaddle collaboration names PP-YOLOE detection, EfficientNet classification and HRNet segmentation. Those examples do not validate every Series3NX or Series4 configuration, and current individual Series4 product pages label the products unavailable, so Series4 should be treated as historical unless sales confirms otherwise.

This layer explains apparent family resemblances. It does not create artifact portability. Two SoCs containing related Vivante or Ethos IP can have different memory maps, drivers, compiler versions, operator sets and vendor runtimes.

What you actually deploy: the artifact lock-in table

The last file in the compiler pipeline is where apparent ONNX portability ends. Names and extensions change over time, but the practical rule is stable: keep the original model and calibration data because the native artifact usually cannot be moved to another NPU generation or reliably rebuilt with another SDK release.

Vendor stackCommon compiler inputWhat reaches the deviceWhat it is bound to
Amlogic AMLNNONNX, TFLite, TorchScript or PT2.adla; legacy A311D uses .nbTarget ID, ADLA generation, AMLNN compiler build, NNSDK2/runtime, ADLA driver and BSP
Rockchip RKNNONNX, TFLite or framework export.rknnRKNN toolkit/runtime version and Rockchip NPU family
HailoONNX or TensorFlow through an intermediate HAR.hefHailo architecture and compiler/runtime generation
Axelera VoyagerONNX or supported zoo definition.axm in alpha Pipeline Builder, .axe pipeline package; classic .axmodel plus manifestVoyager API generation and Metis target configuration
DEEPX DXNNONNX and supported framework paths.dxnnDX-COM/DX-RT version and DX-M target
Qualcomm QAIRT/QNN or SNPEONNX, TFLite or framework modelQNN model library/context binary, or SNPE .dlcHTP architecture, SoC, backend and runtime version
MediaTek NeuroPilotConverted/quantized TFLite.dla for offline Neuron Runtime; .tflite for delegated modeNP/MDLA generation, OS image, compiler and runtime
TI TIDLONNX or TFLiteImported-artifacts directory plus application model filesTIDL tools/runtime and processor generation
NXP eIQUsually TFLite or ONNXBackend-specific Vela, TIM-VX, Neutron or Ara outputThe selected i.MX/Ara accelerator and BSP, not "NXP" generically
ST Edge AI CoreTFLite, ONNX or supported framework modelGenerated network library/code and firmware assetsMCU/NPU target, memory configuration and tool version
Infineon PSOC EdgeInteger-quantized TFLite/Keras/PyTorch pathVela-optimized TFLite with Ethos-U command stream plus application firmwareE83/E84 variant, Ethos-U configuration, Vela, ModusToolbox and BSP; NNLite is a separate target
Renesas DRP-AIONNX/TFLite via TVM or TranslatorRuntime-model directory containing multiple binary data filesRZ/V device and translator/runtime generation
Sony IMX500Quantized/packed network from Edge-MDT flow.rpk packageSensor firmware, memory budget and packer version
SynapticsTFLite/ONNX and supported imports.synap on SyNAP; .vmfb on TorqSL16xx versus SL261x software/hardware generation
Ambarella CVflowPartner toolchain inputPartner-only deployment package; exact artifact contract is not publicly documentedConfirm the CVflow generation, SDK/BSP and camera pipeline through the Cooper Developer Zone
D-Robotics BPUONNX and toolchain-supported graphX5 .bin; current rdk_s .hbm; X3 uses an older platform flowBPU generation, repository branch and matching OpenExplorer/runtime release
AXERA Pulsar2ONNX.axmodelAX chip target, Pulsar2 and AXEngine version
SOPHGO TPU-MLIRONNX, TFLite, TorchScript and others.bmodel on BM, .cvimodel on CV18xxBM/CV chip target and SOPHON runtime generation
Huawei Ascend CANNONNX or supported framework graph.omAscend chip, CANN/ATC and device firmware/driver
Cambricon MagicMindFramework graph or exchange modelSerialized MagicMind engine/modelMLU architecture and Neuware/MagicMind version
Canaan nncaseONNX or TFLite.kmodelKPU generation and matching nncase target plug-in
Mobilint qbONNX and supported framework models.mxqREGULUS/ARIES target and qb compiler/runtime
Rebellions RBLNPyTorch 2 or TensorFlow-supported graph.rblnATOM generation and RBLN compiler/runtime
FuriosaAIONNX/TFLite and supported framework pathWarboy .enf; RNGD .fxbCompletely different accelerator and SDK generations
Sunplus SNNFAcuity-supported network.nbSP7350 NPU driver, runtime and BSP
ESWIN ENNPONNX in the documented current EsAAC flow; other frameworks require export/conversion.modelExact EIC7700-family target and ENNP/ESSDK release
Nuvoton / Ethos-UQuantized TFLiteVela-optimized TFLite with Ethos-U custom operationsM55M1 memory plan, Vela and firmware
Himax WiseEye2Fully integer-quantized TFLiteVela-optimized .tflite in model flash plus board firmware such as output.imgHX6538/WE2 configuration, Vela, TFLite Micro, Ethos-U driver and fallback kernels
Analog Devices AI8XPyTorch checkpoint, network YAML and sample inputGenerated device-specific C source/headers, weights and compiled firmwareMAX78000 versus MAX78002 memory/layer limits, synthesis tool and MSDK release
Realtek AmebaPro2Vendor conversion service/input.nbRTL8735B firmware and VoE/NeuralNetwork API release
Telechips EnlightDarknet, TensorFlow, ONNX or PyTorch path.enlight intermediate plus compiled deployment bundleTCC7500 NPU and TC-NN/Enlight versions
T-Head HHBONNX and supported framework graphhhb.bm, generated parameters/code and executableTH1520 CSI-NN2/SHL stack
SigmaStar MI_IPUCurrent SSU9383CM docs expose the MI_IPU runtime; older material uses TensorFlow/Caffe-family inputsOlder .sim converted to sgsimg.img; current public artifact unverifiedExact SigmaStar IPU and SDK generation; legacy compiler compatibility with SSU9383CM is unverified
HiSilicon smart visionONNX and ATC-supported graph.omExact smart-vision SoC and NNN/SVP-NNN stack; not generic Ascend portability
MemryXONNX, TFLite, Keras or TensorFlow.dfp dataflow packageMX hardware configuration and runtime/compiler release
KneronONNX plus toolchain configuration.nef; KL730 NEFv2 can contain .kneDocumented compiler target, chip generation, firmware and PLUS runtime; KL830 compilation is not established by Toolchain 0.33.1
SiMa.ai PaletteSupported framework or ONNX modelModelSDK compiled .tar.gz, documented executable ELF, or MPK Tool .mpk packageMLSoC/Modalix target and Palette release
NVIDIA TensorRTONNX or network definitionSerialized engine, commonly .engine or .planGPU/DLA architecture, TensorRT, CUDA and often device
AMD Vitis AIQuantized ONNX/framework graphLegacy .xmodel; Gen1 NPU snapshot/subgraphs; Gen2 compiled cache directory or production .rai packageExact NPU generation/IP configuration, platform image and Vitis/Vivado/PetaLinux matrix
Microchip VectorBloxFully quantized INT8 TFLite.vnnx, .hex and .ucomp plus matching firmware/FPGA designCoreVectorBlox configuration, PolarFire SoC Video Kit, SDK 3.1, bitstream and memory layout
Intel OpenVINOFramework model or ONNXIR .xml plus .bin, or device-specific compiled cacheIR is reusable; compiled cache is device, driver and OpenVINO specific
Legacy Google Edge TPUFully INT8 TFLiteEdge-TPU-compiled _edgetpu.tfliteEdge TPU compiler/runtime and supported quantized operators; unrelated to the new Coral NPU IP
Google Coral NPU IPC/C++ examples for the current public IP projectELF examples for RTL/hardware simulation and eventual SoC integration; no public YOLO compiler artifactExact licensed/integrated SoC implementation and evolving open-source toolchain
Lattice sensAISupported network descriptionFPGA bitstream, accelerator configuration and weightsFPGA family, sensAI IP mode and memory implementation
Imagination NC-SDKCaffe, TensorFlow or ONNXCustomer-facing compiled artifact filename is not publicly specifiedLicensed NNA configuration, SoC integration, NC-SDK/DDK and licensee runtime

This table is not merely an extension cheat sheet. The compiler version, target identifier, driver, firmware and BSP are part of the model's reproducible identity. A future LibreYOLO exporter should write them into a machine-readable manifest beside every artifact.

Tool access and lifecycle signals

Hardware availability and compiler access are independent. A board may be easy to buy while its current compiler requires a customer agreement; a public SDK may target silicon that is hard to source. The following are concrete first-party signals, not a universal longevity ranking.

PlatformDocumented signalWhat it means in practice
Amlogic ADLAPublic Apache-2.0 repositories and downloadable wheels; compatible board drivers/BSP may still require Amlogic or a board vendorEasy compiler evaluation, but binary redistribution and the full compatibility matrix should be confirmed
MediaTek GenioPublic IoT AI Hub and Yocto guides; NP8 all-in-one tools and Android documents are marked direct-customer/NDAModel evidence is auditable, while bring-your-own-model automation may require a partnership
HailoPublic model zoo and runtime; Dataflow Compiler distributed through the Developer Zone; public hailo-camera-apps was archived May 3, 2026Broad model discovery is open, but reproducible compilation needs the correct account/version; archive status does not establish Hailo-15 EOL and its current vision-processor application package should be confirmed
Axelera AIPublic documentation and public Voyager SDK; customer support requires an accountA rare self-service accelerator SDK, while product support and procurement remain commercial
Qualcomm DragonwingThe Product Longevity Program lists IQ-9075 as Sampling with longevity through 2038 and QCS6490 through July 2036, while the IQ-9075 product page says ActiveStrong industrial planning signal, but the first-party status conflict requires SKU-level confirmation and does not guarantee one AI SDK remains ABI-stable for the whole period
RebellionsThe current support matrix marks CA02 and CA12 EoL, and CA22/CA25 Active; CA21 is absentProcurement and SDK compatibility are card-specific even inside the ATOM family
NVIDIA JetsonOfficial module lifecycle table lists Orin modules through January 2032 and AGX Orin Industrial through July 2033; developer kits have no lifecycle commitmentDesign production systems around modules, not development-kit availability
Sony IMX500 on Raspberry PiThe AI Camera product brief says production until at least January 2028A concrete minimum for that camera module, not for every IMX500 product
Amlogic A311Y3The 2026 launch page advertises a minimum ten-year product-longevity guaranteePromising new-platform signal that still needs contract and SKU details for a product program
Google CoralThe legacy Edge TPU and PyCoral repositories are archived/read-only; the separate Coral NPU IP repository is activeLegacy software maintenance risk, not by itself a formal hardware EOL notice; it says nothing about lifecycle of the unrelated open-source IP project

Why TOPS is not a buying guide

Peak TOPS can be useful within one vendor and architecture, but cross-vendor comparisons are usually invalid unless all of the following match:

  • Numeric precision, including INT8, INT4, FP16 or a mixed mode
  • Whether one multiply-accumulate counts as one or two operations
  • Dense versus structured-sparse arithmetic
  • Input resolution, batch size and model graph
  • Accuracy after quantization
  • NPU-only latency versus the complete application pipeline
  • Memory transfers and host fallback
  • Thermal state and sustained, not burst, operation

A detector is a pipeline: image acquisition, resize/letterbox, normalization, device transfer, neural inference, decode, NMS, coordinate scaling and application logic. Vendors frequently publish only the neural-inference middle. MLPerf Inference Edge is valuable because it defines scenarios, quality targets and measured system-level rules instead of comparing marketing arithmetic in isolation.

For YOLO deployments, the minimum credible benchmark record is:

exact checkpoint + checksum
source and compiled model format
compiler, runtime, driver and firmware versions
target board and clock/power mode
input resolution and batch
precision and calibration dataset
accuracy before and after compilation
warmup and sample count
preprocess, inference and postprocess latency
power measurement boundary

Without those fields, two FPS numbers usually measure different things.

How to choose an NPU for computer vision

Choose in this order, not by the largest number printed on a product page.

  1. Prove graph compatibility. Compile the exact exported model, not a similarly named zoo checkpoint.
  2. Measure accuracy. Validate the compiled artifact on the real task dataset after calibration and quantization.
  3. Measure end to end. Include input conversion, transfers, decode and NMS.
  4. Check software access. Determine whether the compiler is public, registration-only or available only after a commercial agreement.
  5. Freeze the version matrix. Record compiler, runtime, driver, firmware and target identifiers.
  6. Check operating-system reality. Android, Debian, Yocto, Buildroot, RTOS and Windows support are not interchangeable.
  7. Check supply and lifecycle. A technically excellent chip is a poor production choice if modules, drivers or long-term support are uncertain.
  8. Then compare cost, power and performance. Those measurements become meaningful only after the same model is correct on both targets.

Practical starting points by use case

Use casePlatforms worth evaluating firstReason
Raspberry Pi-specific acceleratorHailo-8L/8 and Sony IMX500Official Raspberry Pi products, software images and concrete developer flows
General Linux PCIe, M.2 or USB add-onDEEPX, MemryX, Hailo and Axelera productsObtainable accelerator form factors, subject to host/driver compatibility
Low-cost Linux SBC or cameraRockchip RK3588/RK3576, Amlogic ADLA where a supported board/BSP is obtainable, AXERA, Canaan K230 or Sunplus SP7350Integrated media pipelines and public model/compiler evidence; verify VIM4 V13A+ for the A311D2 NPU
Mobile and high-volume AndroidQualcomm QNN/AI Hub, LiteRT, Apple Core MLLarge deployed base and maintained mobile runtimes
Industrial Linux and roboticsTI TDA4/AM6xA, NXP i.MX, Renesas RZ/V, HailoLong-life embedded products and camera/IO ecosystems
Very small endpoint or MCUSTM32N6x7, Infineon PSOC Edge, Himax WiseEye2, Analog Devices MAX7800x, Nuvoton M55M1, Alif/Arm Ethos-U, GAP9 and SyntiantTight power and memory envelopes with purpose-built tooling and model-size constraints
Configurable FPGA visionMicrochip VectorBlox, Lattice sensAI and AMD Vitis AI targetsFlexible hardware pipelines, but the accelerator configuration, bitstream and model compiler are one versioned system
High-throughput PCIe visionAxelera Metis, Hailo, DEEPX, Mobilint, MemryX, Rebellions or SiMa.aiDedicated accelerator products and multi-stream positioning
China-centered product ecosystemRockchip, AXERA, D-Robotics, SOPHGO, Huawei Ascend, HiSilicon and CambriconStrong regional boards, toolchains and model examples

This table is a research shortlist, not a performance ranking. Procurement, regional availability and the exact model can reverse the order.

What this means for LibreYOLO

The scalable design is not dozens of unrelated exporters. It is one strict interchange contract plus small vendor compiler and runtime adapters:

LibreYOLO model
    -> deterministic static ONNX
    -> VendorCompiler.compile(model, target, calibration, precision)
    -> native artifact + manifest
    -> VendorRuntime.load() / infer()
    -> shared task-specific decode and Results objects

Every native artifact should have a sidecar manifest containing:

  • Vendor, chip target and board target
  • Compiler, runtime, driver and firmware versions
  • Source checkpoint and ONNX checksums
  • Input names, shapes, layout, colorspace, normalization and dtype
  • Quantization precision, scales where exposed, and calibration-data hash
  • Output tensor names, shapes and semantic meaning
  • Detection task, class names and whether decode/NMS is on-chip or host-side
  • Recorded numerical parity and task-accuracy results

LibreYOLO already provides portable ONNX export, quantization tooling, a direct but deliberately narrow RKNN export, and an honest external-compiler Hailo workflow. Under the criteria used in this guide, Amlogic ADLA is a strong next integration candidate: the toolkit is public, the model coverage overlaps LibreYOLO heavily, and the older A311D path can be kept explicitly separate.

The priority after Amlogic should be based on user hardware and the ability to run accuracy validation, not only compiler availability. DEEPX, Sony IMX500, D-Robotics, AXERA, SOPHGO and Qualcomm are technically attractive; TI, NXP, ST and Renesas become especially valuable when industrial partners can provide boards and long-term CI access.

Watchlist: real silicon without a public integration contract

Excluding a company entirely can make a market map misleading. Including it as "supported" can be worse. These vendors sell, sample or have announced relevant silicon, but the public evidence does not yet establish a reproducible current compiler, artifact and named-model path suitable for a LibreYOLO backend.

CompanyWhat is realWhy it remains a watchlist entry
SamsungExynos Auto V920 includes a dual-core NPU advertised at up to 23.1 TOPSSamsung says its Neural SDK is no longer provided to third-party developers; Samsung ONE/Circle does not prove access to the Exynos NPU
NovatekCurrent corporate milestones name Edge AI and Edge Vision/Imaging AI SoCs; a 2020 milestone mentioned MobileNet, SSD and YOLOv3No current public part-number/compiler/artifact/operator/model matrix; historical YOLOv3 is not current SDK proof
NextchipAPACHE5/NVS2900 and current APACHE6/NVS3000 automotive processors with aiMotive aiWare NPU; the APACHE6 page itself currently states both 12 and 8 TOPSaiWare Studio and a runtime exist, but no public artifact, quantization guide, operator list or YOLO matrix was found; the conflicting vendor figures should not be silently resolved
Black Sesame TechnologiesHuashan A1000/A2000 and Wudang C-series automotive/edge AI siliconPublic product information exists, but not a self-service compiler, artifact contract or reproducible model zoo
IngenicT41/T40 camera SoCs with low-bit NPU claims; Magik AI describes PTQ/QAT and graph compilationNo downloadable current compiler, artifact contract or exact vendor YOLO list was located
FullhanMultiple IPC SoCs advertise 0.5 to 2 TOPS NPUs, while the MC6880 NVR family advertises 4 TOPSNo public NN compiler/runtime/framework/model documentation sufficient for independent reproduction
Goke MicroelectronicsA first-party page for GK7606V1/GK7206V1/GK7203V1 smart-camera parts advertises integrated NPU performance, but is served only over legacy HTTP at publication timeNo public converter, runtime contract or named-model matrix; the route is OEM-channel oriented
Chengheng MicroCH37 advertises 64 INT8 TOPS and FP16/FP32/FP64 modes; the company's 2026 timeline says it entered small-batch productionNo public SDK, compiler contract or named-model validation was found; treat it as an early-commercial, non-self-service platform rather than reproducible deployment evidence
Bouffalo LabBL808 includes the BLAI-100 NPU and current official SDK repositories existThe maintained SDK does not expose a current official BLAI conversion/example path; surviving flows are legacy or community evidence

This watchlist is intentionally evidence-driven. A vendor can move into the deployable tables by publishing a current toolchain or evaluation path, a target-specific artifact contract, and at least one named model with an end-to-end recipe.

What was deliberately not claimed

This article does not say that every named model works from every upstream repository. Vendor zoo models are frequently modified, split at different output nodes or paired with custom postprocessing.

It also does not turn these statements into support claims:

  • A compiler imports ONNX.
  • A chip advertises an Android NNAPI driver.
  • A distributor says the board is "YOLO compatible."
  • A community repository runs one fork of one model.
  • A model compiles without an error.
  • A vendor reports NPU-only FPS without an accuracy result.

Google Tensor phone accelerators and numerous OEM-only custom ASICs are also real, but Google does not expose Tensor as a general self-service compiler target comparable to the platforms above. Company existence, an NPU block diagram or Android acceleration is not enough to invent a LibreYOLO integration claim.

Likewise, cloud accelerators such as Google TPU, AWS Inferentia/Trainium, Microsoft Maia and data-center-only AI cards are outside scope. This guide is about hardware a computer-vision developer can plausibly deploy at the edge.

Model licensing is a separate layer. A chip vendor publishing a YOLO conversion recipe does not grant rights to use or redistribute the upstream weights, training code or compiled derivative. Hardware support, SDK licensing and model licensing must each be checked for the intended product.

Research methodology and corrections

For every company, the research looked for four primary-source surfaces:

  1. A current product page or datasheet identifying the silicon.
  2. Compiler and runtime documentation explaining the real deployment path.
  3. A model zoo, compatibility table or end-to-end tutorial naming vision models.
  4. A lifecycle or access signal showing whether developers can actually obtain the tools.

First-party GitHub organizations count as vendor sources. Board-vendor documentation is labeled as ecosystem evidence when the chip company itself does not publish the same material. Third-party performance claims were excluded from the comparison tables.

The completed draft was then re-audited in separate vendor-cohort and cross-table passes. Those checks revalidated target scope, artifact names, precision, model-versus-postprocessor evidence, announced-versus-shipping status, lifecycle labels, benchmark denominators and every entity count. When two first-party pages conflict, this guide exposes the conflict instead of silently selecting the more convenient value.

This market changes quickly. If you work for one of these companies and a chip, SDK, model list or access status is wrong, send LibreYOLO the exact public documentation URL and version. Corrections supported by primary evidence should replace this text; undocumented marketing assertions should not.

FAQ

How many edge AI NPU companies are there?

There is no universal company total because product vendors, IP licensors, smart sensors, GPUs and private automotive ASICs overlap. This non-exhaustive guide tracks 69 companies and platform ecosystems with relevant edge AI chips, accelerator platforms, licensable NPU IP or credible watchlist silicon as of August 2026.

What is an NPU?

A neural processing unit is specialized hardware for neural-network operations such as convolutions and matrix multiplication. Vendor terms include NPU, AIPU, BPU, KPU, DLA, HTP, TPU and MLA, and their compiled models are generally not portable between companies.

Which NPU companies officially support YOLO models?

Amlogic, Hailo, Axelera AI, Rockchip, Qualcomm, MediaTek, DEEPX, D-Robotics, AXERA, SOPHGO, Huawei, Cambricon, Sony IMX500 through Raspberry Pi's official AI Camera repository, STMicroelectronics, Renesas, Lattice, Himax, Microchip and several others have first-party model-zoo entries, tutorials or validated results for at least one YOLO generation. The exact generation, task, chip target and SDK version still matter.

Can any ONNX model run on any NPU?

No. ONNX is an interchange format, not a hardware compatibility guarantee. An NPU compiler must support every operator, tensor shape and data type in the graph. Static shapes, graph cuts, custom operations and CPU fallback are common.

What is the best NPU for YOLO?

There is no universal winner. Hailo, Rockchip, Axelera AI, DEEPX and Amlogic have strong documented YOLO coverage, while Qualcomm has exceptional device reach. Choose using accuracy after quantization, full-pipeline latency, sustained power, price, availability, SDK access and operating-system support.

Why can NPU TOPS numbers not be compared directly?

Vendors may count different precisions, sparse operations, multiply-accumulates and peak-utilization assumptions. TOPS omits memory traffic, unsupported operators, preprocessing, postprocessing and host overhead. Compare the same model, precision, resolution, accuracy and full pipeline.

What is NPU calibration?

Calibration runs representative, usually unlabeled deployment images through the model so the compiler can estimate activation ranges for INT8 or another low-precision format. Random or unrepresentative images may still produce a compiled artifact while damaging task accuracy.

Does LibreYOLO support edge NPUs?

LibreYOLO exports ONNX and several runtime formats, has a direct RKNN compiler path for selected Rockchip models, and documents the external Hailo compilation flow. Every other vendor-native artifact still requires its vendor SDK and should be described as an integration candidate until it is compiled, validated and run on hardware.