Remote site, Ignition Edge with the MQTT Transmission module publishing about 900 tags to the cloud broker over a 4G link. Every time the link drops, and it does a few times a week, there's a ten minute hole in the cloud historian, sometimes fifteen. The Edge box itself stays up the whole time.
I tried QoS at 1 and then 2, no difference, and took the keepalive down to 30 s so it reconnects faster. It does reconnect faster. The hole is still exactly as long as the outage.
Carrier says put a second link in, IT says put a second link in. I'd rather understand why the data doesn't catch up first. Any ideas?
900 tags at what rate? And is the Edge historian actually storing through the outage, or just holding the last value for each tag?
1000 ms scan class, publish on change, maybe 200 changes a second on a busy minute. The local historian has every sample through the outage, I trended it on the Edge box, no gap there at all.
QoS only covers a message the client has already handed to the broker. If the client's disconnected nothing is queued unless the client queues it, and that's store and forward, a Transmitter setting, off by default if I remember right. Turn it on, size the buffer for the longest outage you expect, and remember 200 changes a second is a message every 5 ms, so it adds up fast. Then check the tags go up with the source timestamp from the Edge box, not the broker's arrival time, so the backfill lands in the right place on the trend instead of stacking up at the reconnect. A second WAN link only makes the holes rarer. There's a writeup on this here: https://plctr.com/utilizing-iiot-industrial-internet-of-things-with-plc-systems/
History store and forward on, buffer at a million messages, and the box says that's fine on disk. pulled the antenna for five minutes as a test. Data came through on reconnect, but on the cloud side it all landed at the reconnect time, one big spike, not spread over the five minutes.
Cloud side is stamping on arrival. Make it read the timestamp in the payload.
The cloud ingest was using its own clock, a setting on their side called use message timestamp, off. Turned it on, repeated the antenna test, backfill lands where it belongs.
So it was store and forward off on the Transmitter, everything during the outage was never sent. Enabling it with a buffer sized for 200 changes a second off that 1000 ms class fixed it, plus the timestamp setting on the ingest. Second link's off the table for now.
What I still don't know is why the holes were sometimes longer than the outage, maybe the reconnect backoff. Thanks mira88, redgum.
The longer holes are probably the reconnect backoff, ours does the same thing.