FDE PulseFDE jobs open 448New in 7 days 30Companies hiring 52Remote-friendly 25%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Tools

Siemens, AWS or Azure in the factory: ask about the 101st gateway and the 73rd hour offline first

In a factory, an FDE's first job is not to pick the best platform but to locate the problem: automation code, the equipment data model, or the data path at the edge.

Siemens, AWS or Azure in the factory: ask about the 101st gateway and the 73rd hour offline first
Photo: Homa Appliances / Unsplash

In brief

  • Siemens Industrial Copilot is a family of copilots Siemens develops with Microsoft. The engineering edition uses Azure OpenAI models to write and optimise automation code in TIA Portal; the maintenance edition runs on Senseye on Azure.
  • AWS IoT SiteWise is a managed service that collects, organises and stores equipment data using asset models. Each gateway can connect to at most 100 OPC UA servers.
  • Azure IoT Operations is an edge data plane that runs on Azure Arc-enabled Kubernetes and is built around an MQTT broker. It can run offline for no more than 72 hours.
ShareLinkedInFacebookX
Dot chart on an hours axis from 0 to 96. A dashed line marks the 72-hour offline limit of Azure IoT Operations. A weekend network outage lasting 62 hours stays within the limit. A four-day holiday lasting 96 hours is shown as an orange dot in the zone beyond the threshold.
A WAN outage from 6pm Friday to 8am Monday (62 hours) stays within the offline limit of Azure IoT Operations. A four-day holiday (96 hours) exceeds it. So ask about outage history at the discovery stage. Source: Microsoft documentation on Azure IoT Operations; the scenarios are illustrative examples from the article.

Microsoft’s documentation states that Azure IoT Operations can run offline for at most 72 hours, and that performance may degrade during that period. For a factory with an unreliable connection, that number matters more than any feature on the brochure.

Limits like this are something every FDE working in industry has to ask about sooner or later. Siemens Industrial Copilot, AWS IoT SiteWise and Azure IoT Operations often come up in the same meeting, but they do not do the same job.

Understand where each tool sits and where they overlap, and you will avoid proposing the wrong architecture in your first week on site.

One tool for engineers, two ways to move the data up

At the bottom is the equipment, which communicates over OPC UA or MQTT. Siemens Industrial Copilot sits at the top: an application that helps engineers work faster. SiteWise and Azure IoT Operations overlap. Both have components that run on the shop floor and both move equipment data to the cloud; they differ in emphasis.

Siemens Industrial Copilot AWS IoT SiteWise Azure IoT Operations
Main job AI applications for engineers Collection, modelling, storage; gateway at the edge Edge data plane: ingest, transform, route data
Core concept A copilot for each stage of the value chain Asset model MQTT broker and data flows
Where it runs Tied to TIA Portal; Senseye on Azure Cloud, plus a SiteWise Edge gateway on Linux or Windows Azure Arc-enabled Kubernetes
Limit to ask about early Efficiency figures come from pilots At most 100 OPC UA servers per gateway Offline for at most 72 hours

Siemens Industrial Copilot: AI for automation engineers

Siemens and Microsoft first announced the idea at Hannover Messe 2023. The engineering copilot uses large language models from Azure OpenAI Service to help engineers write and optimise automation code. In April 2024 the copilot was connected to TIA Portal, and Siemens said it would be available on the Siemens Xcelerator marketplace from summer 2024.

In March 2025 Siemens extended it to maintenance with a copilot running on Senseye Predictive Maintenance on Azure. According to Siemens, the first pilots showed an average 25% reduction in reactive maintenance time. The company now presents Industrial Copilot as a product family covering the whole industrial value chain: design, planning, operations and service.

Treat the 25% figure as a hypothesis to test. Before switching on the maintenance copilot, measure the customer’s current incident-handling time for a few weeks. Without a baseline, you will have nothing to compare against at the end of the project.

AWS IoT SiteWise: the real work is asset modelling

AWS defines SiteWise as a managed service for collecting, storing, organising and monitoring data from industrial equipment at scale. Its central concept is the asset model: a declarative structure that describes the raw data and derived metrics of equipment and processes.

Picture a factory with 140 OPC UA servers spread across three production lines. The SiteWise Edge gateway software reads OPC UA data on the shop floor, but each gateway can connect to at most 100 servers. So you need at least two gateways, and you must decide whether to split them by production line or by network location before buying hardware.

The gateway runs on Linux or Windows. Besides OPC UA, data can also come in over MQTT via AWS IoT Core, or through the SDK. Storage costs are controlled with three tiers: hot, warm and cold. The cold tier is customer-managed S3, and the retention policy decides how long data stays in each tier.

Azure IoT Operations: Kubernetes on the shop floor

Microsoft describes Azure IoT Operations as a unified data plane for the edge, running on an Azure Arc-enabled Kubernetes cluster. At its centre is an MQTT broker running at the edge, built for event-driven architectures. Data flows transform the data and send it on to cloud endpoints.

The platform supports MQTT and OPC UA and targets predictive maintenance, energy optimisation and digital inspection. Because it runs on Kubernetes, the question to ask early is: who in the factory will operate that cluster, and has the customer’s OT team ever run Kubernetes themselves?

Now back to the 72-hour threshold. If the WAN drops from 18:00 on Friday to 08:00 on Monday, that is 62 hours, still within the limit. A four-day public holiday, however, runs to 96 hours, well beyond it. So ask for the outage history at the discovery meeting, not at go-live.

What to learn first

Then practise asset modelling: separate raw measurements from derived metrics, and name things so that operators can read them at a glance. If you are aiming at Azure projects, you also need Kubernetes, to the point where you can build and debug a small cluster yourself.

A tip for filtering job ads: map the job description against the layers in this article. If it mentions OPC UA or edge gateways, you will most likely work on the data layer; if it mentions TIA Portal, you will be closer to automation engineers.

On your CV, do not just write “IoT experience”. A line such as “modelled 40 assets, split across two gateways, designed retention across three storage tiers” shows a hiring manager that you have done real work on site.

Reading product documentation takes little effort. Customers pay FDEs because they need someone who knows where the system will break: when the 101st gateway appears, or when the network outage reaches its 73rd hour.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
6 sources
Read next on the roadmap · Stage 3: Applied AIVercel AI SDK: the client's model may change, but the tools you write stayWhen a client has not settled on a model provider, the layer of code between the application and the LLM often decides how long your demo survives.