Hacker Newsnew | past | comments | ask | show | jobs | submit | celrod's commentslogin

I tried it a few times and liked the speed, but often found it ended up looping, i.e. repeating the same token sequence (e.g. the same sequence of 5 paragraphs) over and over again until it hit the max output limit. This doesn't end up happening every session, but does every now and then.

My impression of DSv4.1-flash was very positive aside from this. But that was enough for me to stick with GLM-5.3(-flash), which both gave me consistently great results

I was using a vibe coded bare bones harness. I was wondering if this was normal from DSv4.1-flash, or if its my harnesses fault.


I've had that looping issue with open models too. But never 4.1. I wonder if it's a model + harness combo? But yeah, one loop issue and I'm done with a model forever.

Harness. Especially if a tool call error doesn't say what to do next and the model is not RL'd with that tool, a retry storm is common.

So if you use MCP a lot, simplify the params, be more lenient on validation and rework the errors.

It is quite good with shell.


Yeah, that's what I'd been leaning towards. No mcp, but I'll see if I can reproduce and debug it, since other people don't seem to have that problem as badly as I've experienced it (and the idea of having a nasty bug like that bothers me).

No mcp support. I'll try copying deepseek harness's basic tool call formats as a starting point.


I was just trying deepseek v4.1 flash on OpenRouter. After running reasonably well for a while, I return to the window and see that my scrollback is nothing but this repeated over and over again:

``` OK.

Let me write.

Let me go.

OK.

Let me write the script.

Let me go. ```

I'm not sure to what extant this is a model problem, vs some providers being fairly broken. If I chose a single provider, I could know how to blame and to avoid them. With OpenRouter, I don't know which provider I was on when this happened.


If Q4 takes less than 1.3x as many tokens as bf16 or q8, it could still end up being faster, given how decode tends to be bandwidth bound. The kv cache was still bf16, so a few ops are the same between quants.


Mercor, Tacit Labs, Handshake AI... I suspect companies like these play a big part in model improvements, generating high quality benchmark/task-focused data for training.

However, these do require educated, white collar, workers.


I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.


I'm curious about this because I've had Fable decompile games and help me understand what's going on inside the game itself and it never complained. I'm not sure what it takes to trip the "safety" guards but digging into game code and data files doesn't seem to be a barrier at all. I've used CC to build some personal game mods a few times now. Once for a game with no modding capability explicitly exposed.


Oddly not much to trip it. I once ran “touch AKAMAI.md” and then had Claude run git status and it fell back to Opus as if I had just broken into the pentagon itself.


>and then had Claude run git status and it fell back to Opus as if I had just broken into the pentagon itself.

Maybe it's more like you asked a lieutenant to check the weather forecast for you, and it went away and sent back a sergeant in its place.


> A narrow set of frontier LLM development tasks, such as distributed training infrastructure, ML accelerator design, and kernel development for certain non-standard chips.

https://support.claude.com/en/articles/15363606

My work is mostly on the Nvidia b200, which apparently gets flagged as non-standard.

Opus 5 works, but sometimes I do wonder if it's surreptitiously trying to sabotage the efforts -- possibly deliberately, but more likely by something like Fable's initial launch, which did come with secretly degraded performance when detecting kernel work. Anthropic was open at the time that such a mechanism existed, but disabled it due to backlash. More likely than not, this is just paranoia on my end...


GLM 5.3 is supposedly incredible for kernel engineering. Have you tried it?


I think I'd rather have the model stop once the obvious solutions failed and ask me. It can suggest more creative ideas, but I don't necessarily want it to try implementing them.


A few years ago, on July 4th one of my mom's dogs freaked out, somehow managed to escape, and got hit by a car before my mom found her. She loved that dog, regularly attending nose work competitions with it. One of your pets getting a seizure must be harrowing for both you and the dog.

I don't light fireworks.


My female Weimaraner was terrified of thunderstorms and fireworks. Once, she managed to squeeze herself under my Mini Cooper. I didn't know at the time that panic vests were a thing. The breed is a gun dog ¯\_(ツ)_/¯. I forget how we managed to get her out from under there, but it took a bit. My Akita doesn't care, he checks the balcony, yawns, and back to his toys. But he is a afraid of air balloons and blow up Christmas Santas. Go figure.


When claude goes down a wrong path, I tend to clear context and write a new prompt that helps guide it down the correct path. Whatever thinking or context that led it there has inertia and tends to be sticky, otherwise.

Pretty annoying when it brings those up again later from memory...


I found this, which has some: https://arxiv.org/pdf/2605.28876 TLDR: RTK does not look good according to the author's benchmark.


That list also places Sonnet 4.6 above Opus 4.6, which doesn't match my experience.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: