Indirect Addressing on a Logix Array: When the Subscript Goes Out of Range
A 1756-L83E drops to Program mode on a Type 4, Code 20 and the fault record names a routine and a rung. Go and look at that rung and it is false — false when the controller faulted, and false since the machine started, because the XIC in front of it is the guard bit somebody put there specifically to stop this happening.
The guard works on the instruction you were thinking about. It does not work on the one beside it.
1756-PM004L says it in two sentences on page 48, and they are the two sentences this article exists for: every instruction generates a major fault if the array subscript is out of range, and transitional instructions also generate a major fault even if the rung is false — the controller checks the subscript in those instructions whether the rung is true or not.
So a MOV behind a false XIC is safe. A ONS behind the same XIC is not, and nothing on the screen distinguishes them.

Ten elements means subscripts 0 through 9. The recipe numbers on the operator’s screen start at 1, which is where most of these begin.
What Code 20 actually covers
Two different faults share the code, and knowing that saves you looking in the wrong place.
1756-PM014O spells it out. Type 4 means an instruction in this program caused the fault, and Code 20 means either an array subscript is too large or a .POS or .LEN value of a CONTROL structure is invalid. So the same two numbers arrive from a MOV with a bad index and from a FAL whose control tag has been trampled, and the record does not say which.
If the rung it names carries no array subscript, stop looking for one and read the control tag instead.
The fault-code table is not in that manual any more.
1756-PM014O carries the procedure for handling major faults, the GSV on MAJORFAULTRECORD, and an example fault routine, but where the table used to be there is now a pointer to the Logix 5000 Controller Fault Codes spreadsheet in the knowledgebase. That is worth knowing before you go hunting a PDF at two in the morning, and it is the same story as the timer fault in TON, TOF and RTO on a false rung, where a negative preset throws Type 4 Code 34 and you are sent to the same spreadsheet.
The guard bit that only guards half the rung
Here is the pattern almost everybody writes, and it is right as far as it goes.

Both rungs are false. The MOV never evaluates its subscript. The ONS has its evaluated anyway.
A LIMIT upstream sets Index_Valid when the recipe number is between 0 and 9, every rung that touches the array is gated on it, and the array is now safe from an operator who types 10. That is good practice and you should keep doing it.
What it misses is that transitional instructions are not conditional in the way the rest of the rung is — the controller checks their operands regardless, because a transitional instruction has to know the rung condition in order to detect a transition, and evaluating the rung condition is how it finds out.
The instructions that behave this way are the ones with memory of the last scan: ONS, OSR and OSF.
Put a ONS storage bit inside an array — Recipe_Load[Recipe_No], which is a perfectly reasonable thing to want when you have ten recipes and want a one-shot per recipe — and the guard bit in front of it buys you nothing at all.
The fix is not a better guard. It is to keep the subscript inside the array before anything can see it.
Clamp the index in one place, near where it arrives, and let everything downstream use the clamped copy. On a 1756-L83E running a ten-recipe machine that is a LIMIT with 0 and 9 around the raw HMI value, a MOV of the raw value into Recipe_Idx when it passes and a MOV of 0 when it does not, and then every subscript in the project reads Recipe_Idx and nothing else. Four instructions, one place, and it ends the whole class of problem, because after that there is no rung anywhere holding a subscript that could be out of range. Do not clamp it in five places, and do not clamp it on the HMI, where the next screen somebody adds will not know the rule. Clamp it once, in the controller, and name the clamped tag so a reviewer can tell at a glance which of the two values a rung is using.
The storage-bit half of that ONS is in one-shots on a ControlLogix.
Six ways an index walks off the end

Only the first of these is visible when you go online and look. The rest were wrong for one scan, months ago, and are correct now.
The HMI recipe number is the obvious one and it is the easiest to fix: the operator’s screen numbers recipes 1 to 10, the array is [0..9], and somebody somewhere has to subtract one. Do it on the PLC side and do it once.
The others are quieter.
A counter accumulator used as an index — Fault_Log[Fault_Ctr.ACC] — works beautifully until the log fills, because nothing in a CTU stops at the array boundary; the .PRE is only a comparison, not a limit, and .ACC keeps climbing past it. An expression index is worse, because both operands can be individually legal: 1756-RM018A prints the example itself, my_list[position+offset] with position at 2 and offset at 5 reaching element 7 of a ten-element array, which is fine, and the same two tags at 6 and 5 reaching element 11, which is not. Checking position and checking offset proves nothing, because the thing to clamp is the sum. A SINT index is a trap of its own, wrapping from 127 to -128 with no warning at all, and an uninitialised DINT is the one that looks safest of the lot, because its zero is a perfectly valid subscript and hides the bug until something finally writes a real number into it.
And then there is Seq.POS, which is the case that is not really about arrays at all.
A sequencer’s CONTROL structure carries .POS and .LEN, and an invalid .POS raises the same Code 20 as a bad subscript. If your step time comes out of Step_Time[Seq.POS] you have both failure modes on one rung, and the fault record will not tell you which one fired. That is the point where you stop reading the array and go and look at what writes to the control tag.
Prescan used to fault and now does not
One piece of history is still worth knowing, because it explains a bug that disappeared without being fixed.
1756-PM014O carries a table by controller revision. On revision 11.x and earlier, an array subscript beyond the range of the array during prescan produced a major fault; on 12.x it tells you to check the firmware release notes; and from revision 13.0 the controller automatically clears any fault caused by an out-of-range subscript during prescan. So a project that faulted every time the key went to Run on an old controller stops doing that on a new one.
Nothing about the underlying index got better.
That matters when you migrate a working project onto a newer controller and it appears to fix itself, because the index is still wrong and will still fault the moment that rung executes for real in Run mode. It also matters in the other direction, and this is the part worth reading a fault routine for. The pattern in 1756-PM014O is a GSV that reads the program’s MajorFaultRecord attribute into a tag of type FAULTRECORD, an EQU against 4 on its Type member, an EQU against 20 on its Code member, two CLR instructions that zero both, and an SSV that writes the zeroed record back — at which point the fault clears and the controller runs again. Written against revision 11.x that was a reasonable way to survive a prescan that was going to fault anyway. Written against a 5580 it is no longer clearing a prescan fault, because there is not one to clear.
Clearing Code 20 in a fault routine is loading the wrong recipe and carrying on.
The dead end: it is not the array size
When a Code 20 comes in, the first thing everybody does is open the tag editor and count.
That is a fair move and about one time in ten it is the answer — somebody resized an array from 20 to 10 without checking what indexes it, or imported a UDT that shrank a member. The other nine times the array is exactly the size it always was, the index is sitting at a perfectly legal 3, and you are looking at a system that is behaving correctly at the moment you are looking at it. Going round again on the array dimension is the classic wasted hour here, and so is stepping through the logic hunting the bad value, because the bad value was there for one scan and is not there now. Neither the tag editor nor the online rung can show you a number that existed for 8 ms last Tuesday.
Two things find it faster than reading.
Trend the index rather than watching it, at the fastest rate your trend allows, and leave it running through a recipe change and a shift change. An index that spikes once an hour is invisible in a watch window and unmistakable on a trend. Or latch it: add a rung that copies the index into a Last_Bad_Index tag whenever a LIMIT on it fails, and the next occurrence tells you the value and the time instead of just the fact. Either beats the tag editor, and the second one survives being left on site.
What to check next
Cross-reference the array tag and look at every subscript on the list, not just the one in the fault record, because the rung that faulted is the rung that happened to execute first and not necessarily the only one at risk. Then find every transitional instruction whose operands contain a subscript, since those are the ones your guard bits are not protecting, and either clamp the index or move the storage bit out of the array. If the index comes from a recipe selection, the UDT and array layout that sits behind it is worked through in moving a recipe into the PLC, and the older introduction to the technique itself, with a worked barcode example, is in PLC indirect addressing.