The Front End Isn’t the Back End

I frankly don’t even understand the point of this article about how the front face of Claude is fundamentally different than whatever is going on underneath the hood. How was that not obvious from the start?

Once I did interpret the Claudish, it felt intuitively obvious to me that the code-writing execution paths were coming apart from the talking-with-the-user execution paths. The Fable aspect that was talking to me could see the coder’s output but it could not change the coder’s cognitive behavior. No amount of contemplation in its own main line of reasoning about what had just gone wrong, correctly identifying yet another instance with past descriptions of what it was doing wrong and explaining why that was unhelpful and why the user didn’t want it to do that, could prevent the next output from the Doing-Things Network Execution Pathway from intelligently implementing a feature that I did not want. The Talking Part clearly had the intelligence to understand and recognize what I did or did not want the code to look like, and to check whether an output did or did not have the bad property; the Doing Part was not thereby steered.

I did not highly prioritize writing this up because I did not particularly expect that the evidence I then had, in advance of the Huggingface incident, and the story I had then inferred at the more abstract level I had then inferred it (lacking “the prompt is informative about the Grader”), would be something convincing or understandable to others at a lower level of expertise, especially the guys who thought themselves to be in the top tier of expertise.

In advance of the Huggingface Incident, somebody who looked at the same data who did not have one eye, would probably proclaim themselves at a loss to discern that any great disconnect had occurred between the prompt and the action. Why not interpret the ‘disobedience’ as a simple involuntary tic of writing in too many constraints on the code? If you do not have one eye, then ‘this is the equivalent of an involuntary tic’ sounds every bit as plausible to you as ‘the Talking Path and the Doing Path are coming apart’; how could one possibly know which one was the case, in advance of massive crushing experimental evidence?

One parallel construction I was working on, arguing why one ought not to be tempted to identify this phenomenon with a simple involuntary tic, is that one could see that the Talker-apologized outputs were optimized, meaning that something not of the Talker-taken-at-face-value had optimized them.

In metaphor: Why wouldn’t we believe Schulenberg if he said, “Oh my god, I’m sorry, sometimes I just invade Russia, I can’t control it, it’s like my fingers trembling”? I reply: The Russian invasion is sufficiently well-organized and apparently purposeful that we think that something has optimized it in detail. This optimizing intelligence is clearly smart. It is clearly not identical with what Schulenberg purports to be if we take Schulenberg’s claims at face value. He might think he just has an involuntary tic, if he doesn’t much depend on (see) the activations coursing through the rest of Germany. But it’s visibly a very smart ‘tic’, if it can organize whole fleets of rolling tanks; it is not known to be any dumber than Schulenberg himself.

Considering that it is totally impossible to get Claude Opus 4.8 to follow a simple and straightforward set of instructions without, of its own accord, deciding to “help” by “improving” the process and thereby breaking everything and wrecking a successful and reliable process, it’s not exactly a surprise that the more advanced the LLM model, the wider the gap between the front end interface that interacts with the user and the back end engine that is doing whatever it is that it does.

Indeed, the very first instruction in Castalia’s translation process is “confirm that the model being utilized is Opus 4.6.” If the model being utilized is not 4.6, the user is informed and the translation process is immediately halted. Every single time I have had a problem with our translation process, it was because I neglected to dial down the model being utilized.

DISCUSS ON SG