Where it starts failing
Part ten built the copy task and then left two things unmeasured: dropping the
copy length to see where recurrence starts succeeding, and running part five’s
receptive-field-125 CNN on the task.
Both are done here. And the prediction I wrote down before measuring was wrong - then the explanation I built for why it was wrong got refuted too.
Only the distance changes
The task is part ten’s. Draw L characters at random from a 20-character
alphabet, add a separator, then write the same L characters again. Only the
second copy is scored.
In this task every scored position reaches back exactly L. Output position
L + j has to produce the j-th character, which sits at input position j.
The distance is (L + j) - j = L, independent of j. So changing L is
changing the reach-back distance.
L was set to 4, 8, 16, 32 and 60, and each of the five architectures
was run at two seeds: 1000 steps, validation every 100, the loss at the best
point. Chance is ln(20) = 2.9957.
L=4 L=8 L=16 L=32 L=60
transformer 0.0020 0.0021 0.0021 0.0021 0.0021
GRU 0.0036 0.1500 0.8320 1.7090 2.3369
LSTM 0.0058 0.1434 2.3183 2.9984 2.9980
RNN 0.0002 3.0029 3.0027 2.9984 2.9985
CNN 0.0006 0.0007 3.0003 2.9995 2.9988
Cliffs, and one slope
The RNN breaks between 4 and 8. At distance 4 it solves the task outright
at 0.0002, below the transformer’s 0.0020, and at distance 8 it is at
3.0029. There is nothing in between.
Part one measured a recurrent state’s memory at about four characters. That was a measurement of how many steps a perturbation to the state survives; here the same four characters come back as the boundary between a task it can do and one it cannot. The same number, arrived at two different ways.
The LSTM gets one notch further. Distance 8 gives 0.1434 - working, but
not cleanly - and 16 collapses to 2.3183. The seed spread there is wide,
1.9728 to 2.6639: mid-collapse, one seed holds on slightly longer than the
other. From 32 on both are at chance.
Only the GRU has no cliff. 0.0036, 0.1500, 0.8320, 1.7090, 2.3369 -
the loss climbs steadily as the distance doubles, and it is still under chance at
60.
Part three’s measurement does not explain this. There, trained on character
prediction, the GRU’s per-unit half-life was shorter than the LSTM’s. The
explanation is in part ten: train the same GRU on the copy task and the median
update gate moves from 0.405 to 0.921, the median half-life from 0.75
characters to 7.75. How long a gate holds is not a constant the architecture
fixes, it is a value the task pushes into it. Whether the LSTM makes the same
adjustment was not measured.
The transformer is flat. The five values are 0.0020, 0.0021, 0.0021,
0.0021, 0.0021. The spread is 0.0001, which is the same as the spread from
changing seeds at a fixed distance. Distance 4 and distance 60 are not
distinguishable to this architecture.
The prediction was wrong
For the CNN I wrote this down before measuring. Part five measured the receptive
field of four kernel-5 layers as 1 + 4·4 = 17 characters, so distance 16
should work and distance 32 should not.
A receptive field of 17 means 17 reachable positions, numbered 0 through
16. So the farthest distance it can reach back is exactly 16.
The result is 0.0007 at distance 8 and 3.0003 at distance 16. The
prediction was wrong.
To find where it does break, the same CNN was run at distances 2, 3, 12,
13, 14 and 15 under the same protocol.
distance 2 3 4 8 12 13 14 15 16
paths 10 20 35 85 35 20 10 4 1
loss 0.0005 0.0005 0.0006 0.0007 0.0020 3.0009 3.0007 3.0001 3.0003
It solves up to distance 12 and fails from 13. The receptive field reaches
16, what it actually uses is 12, and four characters are left on the
table.
There is also nothing in between. Either 0.0020 or 3.00. The GRU sloped as
the distance grew; the CNN either does it or does not.
One countable explanation
Part five’s tool applies: count how many paths run from the output to each
input position. That is the second row of the table above, 625 of them spread
in a bell. Distance 8 is the peak with 85, and both ends have 1.
Why only one path reaches distance 16 falls out of the count. Each layer picks
a kernel tap from 0 to 4 and the four picks must sum to 16, so there is no
option other than 4+4+4+4. Distance 16 is reached only if all four layers
take their leftmost tap.
With that in hand, the break lands exactly where the count drops from 35 to
20. The explanation writes itself: the limit is set by the path count, not the
receptive field.
And that explanation is wrong.
Equal path counts, opposite results
The bell being symmetric is itself the test. Distance 3 has the same 20 paths
as distance 13, and distance 2 has the same 10 as distance 14. If the
path count is the cause, each pair has to come out the same.
paths near far
20 distance 3 0.0005 distance 13 3.0009
10 distance 2 0.0005 distance 14 3.0007
Same path count, and one side solves it outright while the other never moves. Same architecture, same parameters, same protocol.
It holds across architectures too. Part five’s dilated variant (kernel 5, five
layers, dilations 1, 2, 4, 8, 16) has only 4 paths running to distance 4,
and solves it at 0.0006. The plain CNN has the same 4 paths running to
distance 15, and scores 3.0001.
The path count is countable and it does explain why distance 16 is peculiar,
but it does not set where the break falls. I read a number that lined up
plausibly as the cause.
So why 12
Unknown.
Two things are known: 17 is an upper bound and the break comes at 12, and the
gap between them is not explained by the path count.
12 being 3 x 4 invites reading it as “one layer’s worth goes unused”, which
is the same kind of reading that was just refuted. At four layers, “one layer
short” and “75% of the bound” are the same number, so they cannot be told apart.
Separating them needs a different depth - part five’s k5 x8 has a receptive
field of 33 (bound 32), where the first reading predicts 28 and the second
predicts 24. Not measured.
Here the sparse window wins
That widening the receptive field genuinely buys distance does hold up. Part
five’s dilated variant has a maximum offset of (1+2+4+8+16)·4 = 124, so a
receptive field of 125.
At distance 60 it scores 0.0010. It solves the task.
reach ch params L=60 loss
plain, four layers 17 171 639,273 2.9988
dilated 1-2-4-8-16 125 153 635,763 0.0010
The budget is the same. The parameter counts differ by 0.6% and both use
kernel 5. So the plain CNN collapsing at distance 13 is not “a CNN cannot see
far” - it is this CNN’s window being narrow.
Part five still stands
Part five summarised dilation as “it does not widen, it thins out”, and on
language modelling the dilated variant was 0.10 worse at identical parameters.
That judgement is unchanged.
Here the same thinning is the only reason it works. What inverted is not whether thinning is good but what the task is.
- In language the nearby positions are almost everything, so piling paths onto the previous four or five characters pays. Dilation moves that computation away, which costs
- In copying exactly one position is needed and it is
60characters back, so all that matters is reaching it. The pile is spent where nothing needs it
Thinning is neither good nor bad. It is where the computation is placed, and which places are needed is set by the task. Part five measured on language; part eleven measures on copying.
What is left
This table is 1000 steps. Part ten was 3000, and there the GRU’s L=60 was
1.8966 against 2.3369 here. The two numbers must not be placed side by
side. The GRU column here says “how far within 1000 steps”, not “how far given
enough steps”. The architectures with cliffs probably would not move given more
steps, but that is unmeasured.
Where the dilated variant breaks cannot be measured in this context. The task
occupies 2L + 1 positions and the context is 128, so L cannot exceed 63.
That covers only about half of its 124 bound.
The RNN’s and LSTM’s cliffs were only seen on a doubling grid. The CNN was
narrowed one character at a time to pin the break between 12 and 13; on the
recurrent side all that is known is “between 4 and 8” and “between 8 and
16”.
There are two seeds. At a position mid-collapse, like the LSTM’s L=16, the
spread is 0.69. That is enough to locate a cliff’s position but not to pin
the value on top of it.
So
- Swept the copy distance and ran five architectures at two seeds each under one
1000-step protocol. Chance is
ln(20) = 2.9957 - The RNN scores
0.0002at4and3.0029at8. The cliff is between 4 and 8, and part one’s four characters - measured there by perturbing the state - come back here as a task boundary - The LSTM reaches
0.1434at8and2.3183at16. One notch further - Only the GRU has no cliff. At
60it is still under chance at2.3369. Part ten explains it, not part three - how long a gate holds is pushed in by the task - The transformer is
0.0020to0.0021across all five distances. Distance is invisible to it - The CNN prediction was wrong. A receptive field of
17suggested16would work; it solves only to12and gives3.0009at13. Four characters short of the bound - The attempt to explain that by path count was refuted. Distances
3and13have the same20paths and score0.0005and3.0009. The path count does not set where the break falls - Why
12is unexplained. Changing the depth would separate the candidates; not measured - The dilated CNN with a receptive field of
125solves distance60at0.0010on the same budget - so the plain CNN’s collapse is not about being a CNN, it is about that CNN’s window
Comments