Palletiser cell, Beckhoff CX5140 on TwinCAT 3.1 build 4024, an EtherCAT line of 22 slaves. EK1100 couplers, four EL7211 servo terminals, an EL6751 CANopen master on the end. PLC task 1 ms, EtherCAT task the same.
Since last week, a few times a shift, slaves 14 to 22 flip to SAFEOP and the axes on 15 to 18 trip out with a following error. WcState on those slaves goes to 1 while it's happening and the master's Lost Frames counter climbs by a few hundred each time. Slaves 1 to 13 sit there in OP through all of it.
I've tried the task cycle at 2 ms, no change, and I reloaded the config, same again. I guess it could be the master overloaded, but CPU on the Real-Time tab is only 30 percent. Any thoughts?
Slaves 1 to 13 fine, 14 onward bad. Something sits between 13 and 14. How long is the cable between those two, and what sort of cable is it?
About 8 m, from the last EK1100 in the main panel to the coupler in the gripper panel. Standard EtherCAT cable, the green stuff, and it runs through the cable carrier to the gripper. Everything from 14 onward is in that gripper panel, which moves.
A moving cable and a clean cut at slave 14. That's frame loss on one segment, it isn't the cycle time and it isn't firmware. Go and look at where the frames die. EtherCAT master, Online tab, the slave list has CRC columns per port, two counts per box for the two ports. A frame that dies on the link between 13 and 14 counts on the receiving port of slave 14. Everything downstream never sees that frame, working counter wrong, SAFEOP. The ones upstream had already processed it, so they stay OP.
Reset the counters, run a shift, then post whichever slave has the non-zero number. A screenshot of that column will do.
Reset the counters. One shift later, CRC on slave 14 port A is 312. Slave 13 port B zero, everything else zero. I've replaced the 8 m cable with a spare, same green standard stuff, and maybe that was the mistake. Two hours in and its at 40 already. Better. Not fixed.
The spare's the same cable type, so you've bought yourself the same problem a few weeks later. Standard EtherCAT cable isn't rated for a cable carrier. The copper work hardens and the pair breaks strand by strand, and it'll still pass a tester on the bench afterwards. You want the highly flexible drag chain rated version, Beckhoff sell one, ask for the flex rated ZK cable and don't quote me on the digits. One straight run through the carrier, no coupling inside it, because a joint sitting in a section that moves will have that counter climbing again inside a month. Then 14 A stays at zero. There's a writeup covering the diagnostics side here: https://plctr.com/utilizing-ethercat-in-high-performance-plc-systems/
Drag chain rated cable in, one run, no joints in the carrier. CRC on 14 A stayed at zero over two days, no SAFEOP, no lost frames, no following error on the four axes.
So it was the cable in the carrier to the gripper. It work hardened and dropped frames, and every slave past it went SAFEOP while everything upstream carried on fine. The port CRC counters on the master Online tab found the link, and only that one cable got changed, for a flex rated one. EL7211 firmware untouched, the vendor was guessing at that.
Annoying part is the old cable measures fine on the bench with a tester, so a tester would never have caught it. Thanks tobik23, dieter_k.
Same on a gantry once. Tester said fine for two weeks.