Hi all
Cells 3 to 6 went down together yesterday afternoon. They all hang off one Stratix 5700 in the MCC room. HMIs went grey, each cell has a CompactLogix 1769 L33ER and every one of them lost its remote I/O, 16#0204 across the board. Every port light on that switch was flashing at once.
There was a contractor in the room running a new drop for a label printer, and he'd patched a cable between two ports on the same Stratix to test the run. We pulled the patch out, checked each cell, and they were all back inside a minute.
What I can't work out is why one test cable would do that. I'm half thinking the switch is faulty and the extra cable tipped it over, or maybe the printer PC dragged something in with it. Should we be replacing the Stratix? Does one patch lead between two ports on the same switch really take six cells out like that?
Not a virus. A cable from a switch back to itself is a loop, and the switch is supposed to block that on its own. Is spanning tree turned on?
OG
Had to look up what that even was. Logged into the Stratix web page, under Configure there's an STP page and it says Spanning Tree Mode: Disabled. Screenshot attached.
There's a note in the switch description field saying REP ring too, but I can't see a ring anywhere, every cell is a single cable.
A broadcast comes in one port, goes out the patched port, comes back in the other, and with spanning tree off there's nothing to stop it, so within seconds the switch is pushing the same ARPs and the rest out of every port at line rate. Everything on it drowns, which is your 16#0204 on all four cells. Somebody switched STP off, usually during a REP commissioning that never got finished, and that note fits.
Turn it back on. Configure, STP, Rapid PVST or whatever the other switches are running. Then storm control on the edge ports, so a broadcast above a few percent gets dropped. Do it during a stop. There's a writeup on the switch side here: https://plctr.com/advanced-plc-networking-and-communication-protocols/
Turned STP on at lunch with the cells stopped. Ports sat amber for about 30 s then went green, cells all back with no 16#0204. Storm control I found under the port settings, put broadcast at 10%, though I'm not sure that's the right number.
One new thing. Cell 5's HMI now takes about 40 s to come back after a power cycle, and before it was maybe 5.
That's the port sitting in listening and learning before it forwards. Set the ports facing a PLC or an HMI to portfast, on the Stratix the Automation Device smartport role does it for you. 10% is fine for a start.
OG
Smartport role on the eight edge ports and the HMI is back to 5 s. Ran the loop test again on purpose this morning with cell 3 in manual, one of the two patched ports went amber and stayed there, and everything else kept running.
So spanning tree had been off on that Stratix since a REP commissioning nobody finished, and the contractor's test patch made a loop that turned into a broadcast storm. STP back on, storm control at 10% on the edge ports and smartport roles on the device ports sorted it. No 16#0204 since.
I still don't know who turned it off or when, and there are three more Stratix in the plant I haven't been near. Thanks billr and petek.
Retitled for search, it was "four cells down". Marking solved.