# Edge AI in factories and shops: choosing between TensorRT, LiteRT and ExecuTorch

> A model that runs well in the cloud can still fail next to a conveyor belt with a weak connection. FDEs need to plan for the device, the power supply and the rollout before they think about the model.

Bản gốc: https://fdetimes.net/en/guides/edge-ai-tensorrt-litert-executorch/

You arrive at a factory with a defect-detection model that scored well in the cloud. The floor manager walks you to the line and points at a camera mounted above the conveyor belt. The Wi-Fi is patchy, the electrical cabinet is hot to the touch, and he asks just one question: "Can it run right here?"

Plenty of good engineers struggle with that question, because they have only ever deployed to servers. Taking a model into the field is a separate skill. You need to understand the physical constraints of the site, pick the right runtime and plan for dozens of locations, each with its own hardware.

This is also where an FDE adds value that the model team back at the office cannot. They hand you a checkpoint file. Your job is to turn it into a system that runs where the customer actually manufactures or sells.

## What problem does edge AI solve?

NVIDIA defines edge AI as deploying AI applications on devices out in the physical world rather than in a cloud data centre. IBM is more specific: algorithms and models run directly on local edge devices. The definition is simple. The hard part is knowing when a customer genuinely needs it.

There are usually three reasons. The first is speed: according to IBM, when all processing happens on the device, users get faster responses. The other two are bandwidth and privacy, since NVIDIA describes a model in which only the analysis results are pushed to the cloud while the raw data stays on site.

Two examples NVIDIA gives are close to what FDEs often encounter. In a factory, sensors on equipment detect faults and alert managers when a machine needs repair. In retail, businesses want customers to be able to order by voice to improve the shopping experience.

## Work out the budget before choosing hardware

Picture a quality-inspection line with 4 cameras. Suppose each image is 200 KB and each camera captures 10 images per second. Sending every image to the cloud means 4 × 200 KB × 10 = 8 MB per second, or roughly 64 Mbit/s of continuous upload for the whole shift.

A factory's connection rarely sustains that reliably. If the model runs next to the camera, each inference only needs to send a JSON record of a few hundred bytes, along the lines of "camera 2, scratch defect, confidence 0.91".

At 40 inferences per second, the link carries a few tens of KB per second instead of 8 MB per second, and images of the customer's products never leave the factory.

Only now does hardware come in. According to NVIDIA, the Jetson Orin Nano delivers up to 67 TOPS in the smallest form factor of the Jetson range, with power modes from 7 W to 25 W. That power range shapes the design directly: you need to ask how much the electrical cabinet can supply and whether the enclosure can dissipate heat at 25 W.

**Điểm mấu chốt:** Measure bandwidth, power and response time on site first, then choose the model and the runtime.

## Separate hardware from runtime before you choose

The most common confusion is treating Jetson as an alternative to LiteRT or ExecuTorch. Jetson is a hardware module, and inference optimisation on it is usually handled by TensorRT, because TensorRT ships with NVIDIA's software stack.

LiteRT and ExecuTorch are other on-device frameworks; whether a particular Jetson generation runs them well is something you must verify yourself before making promises to the customer.

Thinking in two layers like this, you choose the device according to site constraints, then choose the runtime according to the model and the customer's team.

The first runtime is TensorRT, which comes with Jetson hardware. Jetson is a system-on-module (SoM) that has to be plugged into a carrier board to run. NVIDIA provides the JetPack SDK, a complete software bundle for developing and deploying AI applications at the edge, which includes TensorRT for inference optimisation. Quality inspection on a production line is a typical use case for this combination.

The second runtime is LiteRT, the new name for TensorFlow Lite. Google describes LiteRT as an on-device framework built on TFLite, and the old TFLite guides now redirect to it. The workflow has two steps: convert a PyTorch, JAX or TensorFlow model to `.tflite`, then apply post-training quantisation.

LiteRT is not limited to phones; it can also bring lightweight models, such as anomaly detection or sensor fusion, onto microcontrollers and embedded devices.

The third runtime is for teams already comfortable with PyTorch. ExecuTorch is the replacement for PyTorch Mobile. Rather than relying on TorchScript, it uses the PyTorch 2 compiler and export mechanism. According to the PyTorch documentation, ExecuTorch uses far less memory than PyTorch Mobile, and its memory footprint can be adjusted flexibly.

The comparison table accompanying this article sums up the choice: a multi-camera line with an electrical cabinet leans towards Jetson with TensorRT; vibration or temperature sensors on very small devices lean towards LiteRT on a microcontroller; a counter tablet app written by a PyTorch team leans towards ExecuTorch. Each option comes with a question you must put to the customer before committing.

## From checkpoint to device

Step one is to write a "site budget" before touching any code. That page records the maximum response time (for example, the product is only in frame for half a second), the measured upload bandwidth, the available power and the number of locations to be deployed. You will use this same page to agree with the customer on what "it runs" means.

Step two is to choose a runtime using the comparison table, then convert. For LiteRT, that means exporting to `.tflite` and quantising. For ExecuTorch, it means exporting with the PyTorch 2 tooling. For TensorRT on Jetson, it means optimising with TensorRT. Whichever route you take, keep the original model so you have something to compare against.

Step three is to re-measure the compressed model's accuracy on real images captured at the customer's factory, not just on the test set at the office. Step four is to design the data flow so that only results go up to the cloud.

Step five is to plan for every device across every location: IBM warns that scaling distributed AI to many sites runs into data gravity, heterogeneous hardware, scale and resource constraints.

## The most expensive mistakes usually sit outside the model

The most common mistake is benchmarking on a desk with a cooling fan, then installing the device in a sealed electrical cabinet in midsummer. A device with a 7 W to 25 W power range will perform very differently depending on its power mode, so measure in the mode it will actually use in the field.

The second mistake is quantising without re-measuring, and only finding out when the customer notices more false defect alerts.

The third is planning as if everything were still at the demo stage. Buying Jetson modules and forgetting the carrier boards, or writing handover documentation that still uses PyTorch Mobile when ExecuTorch is available, both cost the customer's operations team time.

The fourth is treating 30 shops as identical, when in practice each may use a different generation of tablet.

For developers looking to move into FDE roles, this skill is easier to demonstrate than it might seem. On a CV, don't just write "knows TensorRT". Write a line with numbers: "cut upload from 64 Mbit/s to a few tens of KB/s with on-device inference; accuracy after quantisation dropped by X points".

When reading job descriptions, look for phrases such as "on-prem", "edge" or "customer site": those are the roles that need exactly this skill.

The next time a floor manager asks "Can it run right here?", the best answer is a fully filled-in site budget, plus an accuracy figure measured on their own products.

**Thử ngay tuần này:**

- Take a small image-classification model, convert it to .tflite, apply post-training quantisation, then compare its accuracy against the original on the same 200 images
- Write a one-page 'site budget' for a hypothetical use case: maximum response time, upload bandwidth, available power and number of locations
- Reread an FDE job description that mentions edge or on-prem, and mark the skills from this article you can already demonstrate with a project

## Nguồn

- [What Is Edge AI and How Does It Work? (NVIDIA Blog)](https://blogs.nvidia.com/blog/what-is-edge-ai/)

- [What is Edge AI? (IBM Think)](https://www.ibm.com/think/topics/edge-ai)

- [Jetson Modules, Support, Ecosystem, and Lineup | NVIDIA Developer](https://developer.nvidia.com/embedded/jetson-modules)

- [What Is NVIDIA Jetson? A Beginner's Guide to Powerful Edge AI Modules](https://blog.aetherix.com/nvidia-jetson-beginners-guide/)

- [LiteRT overview (Google AI Edge)](https://developers.google.com/edge/litert/guide)

- [ExecuTorch Overview (PyTorch docs, v0.6)](https://docs.pytorch.org/executorch/0.6/intro-overview.html)
