Yup, but apparently our cyborg cats can only be kittens and the cyborg mice are probably going to be like 4 feet tall. At least according to the US government.
That's the wiki article for the new one. Says "Increasing LHC luminosity involves reduction of the beam size at the collision point, and either the reduction of bunch length and spacing, or significant increase in bunch length and population."
The linked article says the new one will have "between 140 and 200 proton–proton collisions in every bunch crossing, compared to around 60 during the last LHC run." So the "10x luminosity" seems to be composed of ~3x more protons at a time along with presumably a ~3x tighter focused beam.
I had to chuckle a week or so ago when I was bumbling across the internet and landed on an article that had a link to a nutjob page where the title was something like: "CERN found something so disturbing they had to shut down immediately".
The nutjob article (yes, I did read it, LOL) suggested that they had found some universal truth maybe about God or aliens or something, and it scared them so much they just noped out of the science business to prepare themselves for the inevitable revealing of the "truth".
It is truly no wonder that so much of society is so fundamentally fucking stupid when their trusted information sources are so full of obviously false bullshit.
I can't wait until the LHC is back online zipping tiny things around the ring again in a cosmic demolition derby to find the smallest specs of our reality and give them all whimsical names.
I treat these nutjob articles as a subgenre of science fiction and when they're somewhat well written find them very entertaining. The author cosplaying investigative journalist, one part of the audience acting as if it's a real discovery and another part playing debunking and outrage, is all part of the performance and it's hilarious. Check out the flat earth movement, it's great entertainment all around!
> It is truly no wonder that so much of society is so fundamentally fucking stupid
It's not. You got baited into thinking this is the case.
> are so full of obviously false bullshit.
Well, maybe what's less obvious is that there are a lot more people gawking at how stupid other people can be than fools falling for the bullshit you call out.
OpenEvidence is specifically meant to help clinicians make evidence-based decisions in the diagnosis and treatment of patients, not note transcription.
I think you underestimate just how much money is being poured into LLM SEO at the moment. It's real quiet because they don't want to draw attention and countermeasures from the frontier labs, but this is getting huge investment, and they will have a monomaniac focus on juicing product results whereas the attention of the labs necessarily has to be spread out.
Data curation is important and expensive and frontier labs can afford to do it right. Natural data isn't the limitation, we are already literally out of tokens. It doesn't matter how much you poison things it's not going to stop the progress train.
There's a post every other month where some dude who put nonsense information online celebrates because it actually ended up in some frontier models weights.
If it's easy enough that some randos can do it for fun, what do you think happens when there's commercial interest behind it?
Obviously companies are going try nudging AI towards recommending whatever they're selling. It's a logical extension of SEO - and that's a 100 billion USD industry.
Additionally, if I believed myself to be in some sort of spending - err - AI race, I'd try to poison the data sets of my competitors by putting crap out there for others to ingest.
What does it mean, Is it like when somebody used some coding agent to develop a feature and later input prompts and a resulting PR can be used for training by a presumption that final PR was a correct implementation of a prompt?
Yea it’s rejection sampling, so you have an agent, you take a verifiable problem (people use lots of different verification signals but say unit tests etc) and have the agent attempt it K times. You accept the trajectories (all context, tool use etc, the entire log) that are positively verified and use these as training examples.
The trick is to find the examples that are just in between too difficult and too easy for the existing agent, these have the strongest training signals
The question is not whether it has happened or will continue to happen. Of course it will always be a problem to some extent.
Your original claim is that this will be enough of a problem to prevent models from improving in expert level knowledge. I completely disagree with this premise.
If the models fail to improve, it will likely be due to limitations in the transformer architecture rather than poisoned training data.
And even then, I doubt that the transformer is the best architecture we will ever come up with.
Clearly it doesn’t learn or think like a human does, since humans don’t need many gigabytes of text samples to learn to talk, so there is some room for improvement.
That no one has actually solved the underlying problem at all, and the generation of the example LLM has no bearing on the nature of the fundamental problem.
You are totally misunderstanding my argument then. As I said, garbage in garbage out. Your article is just an example of that. It’s pretty obvious that if you train an LLM on bad data, you will get bad output.
What I’m saying is that the AI labs are handling this not by fixing the “garbage out” part, but by minimizing the “garbage in” part.
The fact that all you could come up with was research (not an actual example of poisoning a real training set) from 2025 kind of proves that this isn’t some kind of widespread, unsolvable problem like you seem to be claiming.
I literally just grabbed a random link. I’ve seen dozens of real life examples of poisoning.
The poisoning issue makes it so that no one can use the internet for training anymore, because more and more internet content is poisoned as a side effect - or poisoned intentionally. And .001% of poisoned data is enough to screw things up if included in the training data.
It’s also one reason why Google search results have been getting so much worse - it’s hard to not find a SEO page with subtly (or not so subtly) wrong AI slop on almost every topic you can imagine. Most folks won’t recognize it, but that’s what is going on if you know what to look for.
One other way of putting it is the ouroborus problem - more and more internet content is AI generated, because of people trying to game the system, and they are making it is indistinguishable from real content as possible to get by the AI detection algorithms.
Anyone trying to train on it just ends up eating the shit from another LLM, which poisons it.
Another name for it is ‘model collapse’, which also doesn’t have a known solution yet.
You’ve seen actual model poisoning? Or have you seen a model return the wrong answer due to what it saw in a search result? Or were they hallucinations perhaps? How do you know it’s due to poisoned training data?
And do you even realize how much data 0.001% of the training data for a frontier models is? They’re trained on 10s of trillions of tokens, meaning you’d need hundreds of millions of tokens of poisoned data.
Some of these problems you mention could become real barriers to models improvements, though there are plenty of countermeasures, such as by focusing on high quality data sources like I mentioned before.
We’ve already probably gotten as much as we’re ever going to get from simply scraping more and more unstructured text from the web as a way to improve model performance.
The type of training being done now is around tool use and solving specific types of problems better, which is the type of training data you simply don’t find lying around on the web.
You’re expecting me to know your job? Give me a break.
I’m wondering the same thing. You keep talking of some grand poisoning problem but can’t point to any specific public information except an article saying that it’s possible. As if that was ever in doubt.
Pretty easy to display one thing to verified browsers (just latest few user-agents from the 10ish different mainstream browsers on the 3 main OSes) and another to anything else.
Yes AI scrapers can easily spoof user-agent, but they fall out of date as the browser updates.
Bit harder to catch them in tarpits and then serve nonsense to whoever ever triggered the tarpit.
>Yes AI scrapers can easily spoof user-agent, but they fall out of date as the browser updates.
It’s a hell of a lot easier for a company to ensure that its scrapers all report the latest user agent string than it is to get everyone and their mother to update their browsers in a timely fashion.
Bernie and AOC (which aren't DNC mainstream, but prominent) had just pushed for a moratorium on "AI data centers" with a definition that includes "that are used for the development or operation of AI models at scale" (trivially sidesteppable by "we build this GPU farm to sell to whoever bids for compute" - which is actually true), plus a bunch of fancy extras bundled in like "The government must review and approve AI products before they are released
to ensure that AI products are safe and effective.", while lacking actual definition of "AI" (given that we had "AI" systems since '50s).
Yeah, the bill has a cause - it recognizes some pain points. But then it haphazardly tries to address symptoms instead of underlying issues (environmental regulations, utility pricing, land use, job security), while pushing vaguely defined regulations that allow arbitrary application. As if misdirected measures and poorly defined laws aren't already a giant issue.
The whole point of Congress is to get a bunch of people with different ideas and hash them out. Ideas like this are an input, not an output, of Congress.
Congress did regulate weapons access when they passed the AECA almost 50 years ago. The rules have been refined via ITAR over time. This isn't arbitrary.
The IRA and the CHIPS Act were the "last major thing" the Dems did, and both were far better policy for tech than anything out of this current administration.
The what? More like "the whims of an eighty year old in cognitive decline and those wishing to curry or keep his favor" - quite an expansive definition of "political decision making".
It wasn’t. Biden largely didn’t do much. The trump administration does illegal things that get struck down in courts on a daily basis. We’re all very desensitized to it.
But yes, Biden was old and cognitively not well. But his “whims” didn’t exist much, and they were always fairly reasonable. Trump is the most unreasonable president, most likely in US history. I would even categorize Andrew Jackson as more restrained.
> I think they’d try to get something through Congress to regulate the industry in a rules-based way.
Is that a joke? We're back in a spat with Iran because Obama refused to engage with Congress, as required by our constitution, to enter the USA in any binding deal.
Any AI actions from the next admin is going to be executive yolos.
sadly there's sites dedicated to buying accounts of all sorts (reddit, x, steam, etc...) that use an escrow-type system so both parties have little risk
It's basically super easy and trivial to buy verified accounts for many many platforms
The trouble is getting caught, and you don't list contact info on your HN profile so any potential buyers would have to leave their email address here and tell you to email them and the people running. Steve are an idiot so any email address that appears there is gonna get blacklisted different connected to a Steam account so you might want to list your email address well throwaway that's not attached to your Steam account account. You might want to list that in your profile. Just saying. I already have a Steam account.
But Kubernetes is solving a much messier and more complicated problem than React. There are numerous similar web frameworks to React in different languages that have been created as basically hobby projects.
Of course Kubernetes is going to be way less fun to use. The problem of managing servers and distributed applications at scale is inherently not fun once you get into the nitty gritty details.
Kubernetes has the basic flaw that it has more scalability than 99.99% of companies need and you could serve almost all the market with a system that supports shared data structures (like IBM's Sysplex) and is more opinionated. An architecture which is less scalable could serve almost all of the systems on the planet and would be easier to work with.
I'll grant that there is essential complexity there, but Kube was built by people who didn't have fear of accidental complexity so it has a lot of it. Look at the whole "YAML sucks" thing which is partially a YAML thing (coulda chose something different) and also a function of the system they are trying to configure with YAML.
Kubernetes' YAML problem stems from its CLI tooling, and yes it was an atrocious choice once templating came in and visited horrors like helm on us. Internally, the the k8s api speaks only JSON, and you can already stuff whatever json you like in a yaml file.
Which is probably why Meta has been getting into all kinds of side projects like VR/AR and AI lately. Because there just isn’t that much they can think of in the social media space that’d be worth doing.
Of course, with how mediocrely those side projects have been going, I’m not surprised Meta is turning to layoffs. They seriously over hired and never really found a good use for all those engineers.
Turns out they have no product vision beyond selling personally-identifiable information to advertisers. I hold them quite a bit in contempt so I would relish in their failure.