Hacker Newsnew | past | comments | ask | show | jobs | submit | bitexploder's commentslogin

We have pretty strong indications of when and why things go wrong. Openness vs rigid thinking are an axis that are very strong correlative factors. A lot more nuance but there is a lot we do know about risk factors so it isn’t a complete question mark.

Anecdotally / personally experienced; bad trips happen “when you get stuck”. The #1 thing to do in such a situation is to change the music. The #2 thing to do is, quite literally, go to the bathroom and take a shit.

It’s extremely analogous to anything that would make a normal day into a bad day.


That makes sense. I think I have approached becoming stuck. I reached a kind of extreme disassociated state once where I felt like I didn't exist and part of my brain started to panic. I meditate a lot. I reminded the part of me that was becoming anxious that we are okay because we are somehow still observing ourselves so we must still exist. That was apparently adequate to calm it lol. It is truly weird to negotiate with some part of your brain and see it work in real time. I think it was a particularly sticky day for my DMN and or some particular loop it was hanging on to? It is always hard to say.

“Inadvertently”.

Someone was probably sad when they saw a compiler work for the first time after years of writing assembly. AI is intellectually different, progressing, and has no visible horizon, but in my life technology never did in computing. The point is that technology changes often. I think having a couple of decades in tech makes the shock easier to absorb in many ways. And also harder regarding something so familiar becoming so rapidly irrelevant.

But I think your message is correct. The agents don't do the hard work of making things useful and reliable for humans. There is more to do now, not less, somehow. The universe has changed, but most of the problems we are solving have not.


Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.

Right, imagine if instead they had coined new terminology that was not obvious and it re coined that - this would be close to a smoking gun

Afaict that didn't happen so there's just lots of speculation


One off might not work but how many n off you have to be is probably smaller than you'd guess, because the model does need to fit cases that are rare and would not be represented well in training e.g. esoteric things or very recently documented things.

You can probably game the metrics that models use to weight potential knowledge akin to SEO. Maybe have some bots parrot your data around a bit in some places online, maybe the model picks up on this and sees it as high engagement and promotes it over the correct data.

Maybe there are ways you can coax out the most optimal way to break into the training set out of the model itself.


Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?

That is a better idea. Ingesting your corpus with a lot of traces that have semantic patterns. Semantic steganography that suffixes well to real math and science (and any) topics. <thinking> heh.

"Semantic steganography" is my new favorite search term – thank you for this rabbit hole.

Hah, np, stego in general is really cool :)

I don't usually defend Apple products, but my AirPods Pro 2nd gen have been rock solid and taken insane abuse. I use them working on cars, runs, hiking, rain, shine, sauna. I drop them. They land in puddles. They smack concrete every couple of months. Sample size 1, but they are probably one of my favorite accessories and things I lug around with me. I use them a lot and if not love /really/ like them and nothing else has come close to the frictionless experience of them.

Although that may be in part due to vendor lock in and not sharing some of their connection handling API / tech.


If you are okay with waiting use GLM 5.3 max. It costs more but still cheap. It is slow, but a very strong worker. Still dollars per day (at most) with heavy concurrent agent running. I load up planning and tasks in Opus or Sol, and just have glm flash workers go to town every night. My project has never advanced more smoothly.

Which versions of flash and at what thinking levels? Which chinese flash models and at what thinking levels? What tasks? What completion rates? How was quality evaluated?

- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash

- Medium for Gemini, high for Deepseek.

- Things like find information, then understand something about it, then send a slack message or email etc.

- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini

- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.

Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.


Thanks... I have been trying to figure out some things. Been doing my own evals. Flash 3.8 does burn a lot more tokens on high. Interesting how smart and not smart it is. For personal use almost impossible to justify the cost of 3.8 Flash cost.

Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper.

From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.

All this really needs evals, the token prices tell nothing.


I am surprised at how well DSv4 flash does in the real world vs many benchmarks. You look at Flash 3.8 and it supposedly beats opus 5 and deepseek is far below.. but they were measuring efficiency, whatever that is… Something doesn’t add up for me on the published benches

For what it’s worth, I am very happy with Jellyfin and the *arr suite. It took a bit of agent prodding to get them all playing nicely together and bypassing Cloudflare CAPTCHAs. However, it's pretty sweet when you get it all working.

> bypassing Cloudflare CAPTCHAs

Can you say more about this? I've never encountered this problem in nearly a decade of running the same stack.


With the *arr stack, some of the trackers it uses to search for a torrent will present a captcha when Prowlerr queries it for a search. If you integrate FlareSolverr and set up Prowlerr to submit queries through it for the trackers that have CAPTCHAs, it won't get blocked by them.

Yep. It really isn’t nefarious or abusive. I keep mine slowed down to not scan often to be considerate. Even private trackers can require it now.

Ah - I use Usenet, not Torrents, that'd explain it. Thanks!

It does seem like for new code that might help. There's some really good logic and wisdom in it, but it has to be applied very contextually to the exact problem you are trying to solve. If an agent is navigating a complex codebase, this could definitely send them off on a refactoring rabbit hole. However, if you have them writing some new code, it could prevent their tendency to yak shave and write new things. So I can see some situational uses for this, but it could get out of hand as well.

Sol is my current favorite model to interact with. So much less BS than Opus 5. Fable 5.1 is okay as is Fable 5 but it has Opus like tendencies. Sol is very good at following instructions and remembering them for a session.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: