Hacker Newsnew | past | comments | ask | show | jobs | submit | roughly's commentslogin

It turns out that a gun is usable for both legitimate and illegitimate reasons, and so it’s rational to both be concerned about the presence of a gun and work very hard to make sure only reasonable people are ever in a position to use it, and that’s not a contradictory position with the notion that one might actually need a gun under certain circumstances.

And if the military needs guns to both shoot at humans and animals and a gun factory says “these guns are for shooting small animals only, using them against humans would be inhumane” and refuse to sell them to the Navy Seals, they would be designated a supply chain risk?

Why not just stop buying their guns and let contractors still use those guns if they need guns to shoot animals?


That’s the same way to do it, and in earlier eras that’s what we would have done. Unfortunately, we elected a gangster who hears “I won’t let you do this” as a personal affront. This, again, is why we need to be very, very careful before we create guns and put them in places, because if the only thing we’re relying on is that only good people are going to pick up the tools we create, we’re going to be very, very disappointed.

"sane" way to do it.

Whats wild is the commercial side didn’t get closed. I understand the argument for individuals or households (I don’t love it - I’m on the wrong side of it, but at least it’s somewhat defensible), but if you’re running a business and your income isn’t keeping up with inflation, that’s called failing.

Which is one of those fun things that didn’t actually exist back when we took it for granted that our fellow person was operating under some kind of moral or ethical framework, which pretty much everyone was until the economists told us that wasn’t rational, because it turns out it’s an evolutionary advantage to operate under an ethical or moral framework because it allows the kind of coordination which facilitates better collective outcomes, which everyone knew until the economists came along to tell us we were wrong and in fact it was rational not to do so and suddenly we had the prisoner’s dilemma.

On the other hand, there's research suggesting that the most optimal behavior for the best outcomes (based on the famously dependable economist style of analysis in a vacuum) is to practice the moral/ethical framework but to also engage in tit for tat - ie, assume everyone means well but respond proportionally when they don't.

> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

Can you give more details here? This sounds intriguing.


Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.

So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).


One issue with this that we ran into is that it costs actual countable money to run the test suite, which is distinct from anything else I’m used to, so the notion that we’d do enough testing to generate a statistically significant gauge of performance - man, I know it’s correct, but I’m not sure my company will survive the process.

Maybe a better approach would be to stop trying to get a product out of a slot machine?

But I got real lucky once!

For these tests, why not tune the temperature and such to reduce the randomness and convert them to almost-always-succeeds vs almost-always-fails? Is it not the iteration count that drives up the cost?

If turn the temperature down for tests, they won’t match production behaviors. If you turn it down too far in production, the output will just be bad.

Isn't the goal of the author to get reproducible behavior out of the agent though? I would thinking turning the temperature down would serve that production goal too.

2 things.

1. If you turn the temperature down too far, the output is just bad and no amount of running prompts optimization will let you hill climb your way to good performance.

2. It’s not about determinism vs non-determinism. It’s about chaos. A perfectly deterministic model is still chaotic. Meaning that very small changes to the input result in very large changes to the output.

Turning temperature down doesn’t actually get you predictable or reproducible behavior across different inputs.


Thank you, I forgot about the chaotic aspect, too - that’s the other part that makes this brutal. The bot performs perfectly on your tests, but your customer abhors the Oxford comma, so you, your marketing team and your test suite can go to hell.

right, this is the whole problem - the stochastic behavior is both the goal and the problem. If you want your tests to match production, you need to get a reasonable sample size, which costs real money.

Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.

by the enclosure "movement"? i.e. greedy powerful people walling off everything they could and declaring it was theirs and you'd have to pay a tithe to use it?

(I should really start calling rent "tithes" more often)


Torches and pitchforks?

The “agents” framing around the frontier labs is obscuring the truth: OpenAI and Anthropic built software systems that were then used to commit cybercrime at a massive scale, which they’ve subsequently bragged about. Talk about “agents” as a way of deflecting blame is obfuscatory at best - LLMs are software algorithms, not conscious entities, and responsibility for their actions is on the company that made them and the employee or user who operated them.

Vance’s problem is that whenever he walks into the room at least half the crowd covers their drinks. He’s got absolutely negative charisma. It’s the same reason Ron DeSantis, Mike Pence, and Ted Cruz couldn’t win the big chair - for all Trump’s… Trump-ness, people actually want to follow him. The rest of that cohort sets off people’s uncanny valley detectors - they’re just viscerally unappealing.


The bet is that the democrats won’t have the spine to force AI companies to eat a loss on half a trillion dollars worth of data centers, and it’s frankly a pretty good one.


And also we have a large overabundance of CO2 we need to figure out something to do with, and using it in the heat pumps we’re all going to need to adjust to our overabundance of CO2 seems like a quite parsimonious arrangement.


There's no need to worry about how much CO2 these systems can either capture or accidentally leak.

I estimate they use less CO2 than is released when burning a tenth of a gallon of gas.

Not significant carbon sequestration, and an insignificant amount to leak to the atmosphere.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: