Hacker Newsnew | past | comments | ask | show | jobs | submit | pera's commentslogin

Everything you say can and will be trained against you

This is the scary part - your most novel thoughts and breakthrough ideas being slurped up and regurgitated as if they were the AI’s creativity

Not only did they steal everything from humanity’s knowledge, the theft continues as now we are all hooked up to the machine


It's not at all scary. I know some of my ideas are poorly remembered copies of other people's work. Whenever I'm trying to build something, I spend time going through technical journals on the topic to see who invented it first and what they discovered that I haven't figured out yet. It's amazing how hours in the library save you days of beating your head against the wall.

I suggest looking at the past history of IP disputes. Humans have been "slurping up and regurgitating ideas" for a very long time. There are lots of examples of parallel creation, rediscovering old ideas independently, telling an idea to the wrong person, and having them claim credit for it.

- Newton/Leibniz clash over who invented calculus. - Niccolò Tartaglia vs. Gerolamo Cardano clash over the formula used to solve cubic equations. This was also an independent rediscovery, as Scipione del Ferro discovered and published the formula earlier. - There are multiple literary works in print, music, and film that have competing claims. - Meccano versus Erector Set: developed about 20 years apart in England and the United States. Unclear if it's independent invention or copied. US developer Alfred Carlton Gilbert claims he was inspired by steel girder construction of infrastructure.

also https://community.thriveglobal.com/10-famous-inventions-that...


> Everything you say can and will be trained against you

So, experts are incentivized to seed LLM data with false-leads to confound it. Already, garbage is being published on arxiv and elsewhere, and many sloppy code-repos too hastening the process. Expert inputs will be in more demand to un-shittify.


To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.

Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.


AGI would produce novel treatments for diseases at rates equivalent to what a human can do today.

Which is to say, not that fast.


Whether AI is AGI does not depend on the speed at which it operates/thinks. Clearly all the theoretical work done by AGI will be done orders of magnitude quicker than humans can do it.

It is an open question to what extent practical experimentation/work will be a bottleneck for the theoretical work. It stands to reason that it is improbable that it will be the bottleneck for 100% of the speed of treatment development.


Wouldn’t that be ASI? I.e. surpassing humans by outputting novel treatments at a far greater rate than normal humans?

You don't need super-intelligence to produce at a greater rate when your tasks are parallelizable.

You are describing superintelligence (ASI) not general intelligence (AGI)

I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.

That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.

Just seems very weird to call getting Fields-medal-level results "inference over a bounded search space" and "not actually challenging".

Plagiarizing on a massive scale to generate works which appear to be Fields-medal-level results is not the same thing as inventing new conceptualizations in mathematics.

No matter how bodly they write the headlines, what has happened in mathematics using Large Language Models is very much "inference over a bounded search space" even if those bounds are immense.

For a comparison of true creation of novel conceptualization in mathematics is submit the works of Martin Hairer, one of which is Introduction to Regularity Structures, [https://arxiv.org/pdf/1401.3014] None of the so called, "novel math discoveries" by any LLM is as enlightening and expands the state of the art in math like any of his writings.


If they’re getting results which mathematicians have been trying to do for decades, then I think the “plagiarism” is extremely socially valuable. If they aren’t valuable results, then why were mathematicians being paid to solve them? I don’t buy this “the journey was the insights we got along the way” stuff that disgruntled mathematicians are selling.

It is irrelevant whether in general public understands the value of the insights, and the lesson that mathematics coursework should have made more clear for everyone is that understanding the process that is required to get an 'answer' is where the entire value of a mathematics education exists.

The 'cheat code' approach to math results, a result which no one understands, and which no one can teach has no real value.

The system of payment for publications in order to support math discovery is simply the narrow 'commercial system' applied to supporting foundational science in the absence of a broader civilization level appreciation for the mathematical arts. Looking to history, from the late renaissance through the early 20th century the support for mathematical discovery was more generally understood and supported by institutional level organizations and more generally understood to be important for the progress of scientific progress by the private and public wealth .

This system enabled the development of topology, numerical analysis, complexity, set and group theories. The lapse in this level of support that did not give mathematicians the same protection from front line deployment in WWI brought that era to nearly a close. Reading about 'Nicholas Bourbaki' might lend some deeper appreciation of the effects of the losses from that shift in collective appreciation of foundational math.

The idea that these LLM's are getting results that mathematicians haven't produced demonstrates the shallow understanding of math in modern times, due in part to the limited accessibility of so much of the prior writings of the entire history in mathematics, whether that be due to few surviving copies of some arcane work in a private library collection, or due to a modern fee for access paywall. One example of this condition can be shown with a small excerpt from a work that I am currently composing:

"In 1805, while computing the orbits of the newly discovered asteroids Ceres, Pallas, and Juno from limited observational data, Gauss developed an efficient method for evaluating trigonometric interpolations by recursively decomposing large sums into smaller ones before recombining the results. Because of a steadfast adherence to Gauss' own personal motto, "Pauca sed matura" (Few, but ripe), Gauss never formally published this specific algorithm nor the conclusions of investigations which also laid the foundations of non-Euclidean geometry. These methods remained hidden in his notes under a manuscript titled Theoria Interpolationis Methodo Nova Tractata which was published in 1866, 11 years after his death, and the Fast Fourier Transform-equivalent approach within it remained largely unnoticed until the twentieth century, when James Cooley and John Tukey independently rediscovered the same computational strategy. His discovery was seventeen years before Joseph Fourier published the original Fourier Transform in his 1822 results on harmonic analysis."

That is to say; Tukey and Cooley were unaware when they discovered FFT that the knowledge had lay hidden in an obscure work for centuries. It should be understood that these 'novel' LLM discoveries are simply the models traversal of the huge corpus of all the maths publications in the training set, collecting and rearranging these techniques into synthetic 'results'. They are attention getting, but they are not new, and the proofs are insufficient to the task of improving the utility of mathematics for humanity.

The 'disgruntled mathematicians' aren't selling anything. They are informing civilization as a whole that having a cheat sheet to the math test only cheats yourself in the end, the same point that math teachers have been making since grade-school. Anyone who doesn't internalize that truth will always need someone else to do the math for them.

To paraphrase Curtis Jackson, ""If you don't know the numbers, you don't know your business."


My original point was: if mathematicians have been seeking a result for decades, and now an llm has got it, then either that is socially valuable, or mathematicians have been wasting public money.

Are you now saying that they weren’t really looking for the result after all, but were looking for a psychological state of insight? What is the point of insight? I thought its point was that it led to results.


> "What is the point of insight? I thought its point was that it led to results."

The 'results' are a substitute product for the real social benefit, which is an availability of mathematically educated and educable society. There has been no 'waste' of public money except when the results of publicly funded research is published by for profit corporations and held from the public's access behind paywalls.

It is unfortunate that one outcome of this situation is apparently a public which perceives that the final 'result' of a math research project as the actual product of value, when it has repeatedly been shown by history that the most significant value is within the multiple alternate potential paths explored by other researchers working toward the same result.

Most of these do not lead to the specific result, and many may instead demonstrate that a specific approach conclusively does not lead to the initial objective result. This also has value for humanity, in many cases leading to new paradigms of thinking about similar or unrelated problems which may later provide foundational insight for new approaches to solving different problems in a manner not yet known or discovered. The value of the system is not enclosed within the single results, but in the collective search and expansion of human consciousness and ingenuity that the search for all of these results entails.

This point circles back to my original. When we cheat on a math test or homework, thinking that the objective is to get the answers right, we are only cheating ourselves out of the learning that would have enabled us to develop the correct answers on not only that one test, but on the unknowable challenges which will later arise that require the new solutions to build upon that learning.

The 'prize money' for solving the biggest hurdles in math is less about those specific problems and their specific answers, but on encouraging many people, who do not get 'first place' and win the prize, but who do work toward it in their own unique ways and in turn provide an uplift to the general capability of humanity to solve hard, as yet undefined problems large and small.

Having an LLM give us the answer is entirely missing the point of the challenge in the first place, and robs humanity of the opportunity to improve its collective capability through making the effort. It also does not find any unique perspectives, which new and unique perspectives have formed the basis for civilization scale improvements since the dawn of time.

A good example of this difference comes from Bruce Schnier, who teaches public policy at the Harvard Kennedy School and the Munk School at the University of Toronto from an article originally published in The Guardian, which compares LLM use in education to having a forklift at the gym. It may be the right tool in a warehouse to get heavy things across the floor quickly, but using one to do your workout is both overkill for the amount of weight you need for arm curls and bench presses, and accomplishes nothing in the development of your strength or cardiovascular health.

If the objective is to program another widget, and you can get it done in a fraction of the time with an LLM, sure, why not. But, if the objective is to expand the corpus of human knowledge, which is the most beneficial and desired outcome of the study of mathematics, then outsourcing that task to electrons through semiconducting silicon is both overkill which results in a 'proof' larger than the entire MatLab code base, and does not accomplish the stated objective. Humanity is not made any more capable through this process.

For a deeper insight into what the measurable benefits for individuals and society, an in depth study of the neurobiology results from the study of complex topics; such as linguistics, mathematics, and music theory might help with the comprehension of the less direct value of the process. Challenge yourself to discover what the effects of later-in-life study of foreign language have on the factors leading to senility and mental decline, and see if a study of mathematical theories has any observable effects on the brain's electrochemical development across different ages. Identify how these kinds of study can have collateral improvements for other fields of study, such as material science, medicine, or applied physics.

So, yes. I am saying that. In addition to looking for the result, because it's cool, I am saying that mathematics is indeed searching for the inherent uplift to humanity which are available through many distributed states of psychological insight that mathematics prizes can serve to inspire. The results are simply one of the points of this search, and if all of these searches are cleared off the board mechanistically, we may lose the greater race against our own potential for growth in the trade-off.


My point was that we won't care about benchmarks anymore because we would see an obvious and completely unprecedent increase in productivity (and I believe it will likely come from the same people who will develope such machine).

The reason most of the conversations are focused on benchmarks is because we are still in the age of weak AI.


Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).

If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.


Yes it is a very serious problem because it confuses a lot of folks with a great deal of power like judges and policymakers.

The first book I ever read on ML (late 90s) dedicated the entire first or second chapter exploring the distinctions between artificial and biological neurons, and even talked a bit about the philosophy of modelling. I still remember thinking back then why would the authors spend so many pages on this but now I believe it was because they understood that a metaphor can be a double-edged sword.


To be fair, the ANN architecture underneath is a misleading thing to be looking at, it's not where the comparison comes from. Though I can't tell if you meant it to be relevant in that way, or just as a general example for the dangerous nature of metaphor.

LLMs are expressly designed to approximate human behavior within the bounds of the written word. The anthropomorphization is no more philosophically problematic than saying differential calculus measures curves.


I meant it in the latter way: a metaphor can be useful as a pedagogical tool to introduce new ideas, and using the source of inspiration for this idea as the metaphor itself makes perfect sense, but unfortunately our brains seem to be prone to assign other properties of the metaphor that don't actually belong to the object of study.

I imagine this happens because we tend to conflate things that are similar, or maybe because it's not entirely clear which characteristics are being mapped in the metaphor?


They spent so many pages discussing it only to show that the mechanism for how ANNs work is different from the mechanism for how biological brains work. It says nothing about whether they can compute the same things.


Save that for your local mom-and-pop store: Microsoft is a multi-billion dollar corporation with enough resources to, at the very least, provide a reliable service for enterprise clients.

The "AI is using a lot of resources" excuse was maybe acceptable last year but not in Q3 2026.


Doubly so given this is hardly a surprise. We've been on this trajectory for at least a couple of years now. They don't get to shrug, point at 10x volume, and act like they've been blindsided.


They had large increase in volume, and trying to move to Azure at the same time. I don't envy them for either work they need to do. But also don't feel pithy because it's Microslop, at the end of the day.


All being said, I have most sympathy towards engineers of all people. They don't get to call the shots, they do as they are told. 99% of the time by an extremely out-of-touch managers.

Managers and executives though? Those I do blame. Surely at this point it should be blindingly obvious their current strategy is not working.


Those engineers also chose to work in US big tech for the ridiculously high salaries that come with it. Having to deal with the problems resulting from that scale is just part of the deal.


The end of the day is now.


Ok but if the same problem will happen with the mom-and-pop store, switching there doesn't help. This isn't a moral dilemma, people just want their stuff to be up.


If you use Azure you’d know they are definitely not able to scale up quickly. Tons of quotas and lack of availability.


They are also at least partially responsible for and profit from the LLM spam.



Yeah it's not just Amazon, but so far Amazon was the only one specifically looking for URLs in source code. Interestingly it ignored URLs in the fake markdown docs.

Another detail: the scraper did not attempt to access the endpoints immediately (as it did for hrefs in htmls) but it did it on the day after, twice.


Was it actually an Amazon bot IP[1] or someone pretending to be on AWS?

1. https://developer.amazon.com/amazonbot/searchbot-ip-addresse...


Yup, already mentioned in a comment below: the IPs are in that list.

They also show up in AbuseIPDB with multiple reports.


Thanks, the Amazonbot IPs that accessed my honeypot are in this other list:

https://developer.amazon.com/amazonbot/searchbot-ip-addresse...

By the way, they used false user-agents.


My concern is more in relation to trying to use an endpoint from non-public source code, it seems negligent to randomly try endpoints like this


Whether it’s negligent isn’t terribly relevant. Is it illegal? It is not, they aren’t bypassing access controls. No different than using Shodan, scanning public IPs, crawling open directories, etc.

It would be different if they were attempting to brute force credentials to access an endpoint, but they aren’t.


> they aren’t bypassing access controls.

Weev went to jail for accessing public api's, https://en.wikipedia.org/wiki/Weev#AT&T_data_breach

> The flaw was part of a publicly-accessible URL, which allowed the group to collect the e-mails without having to break into AT&T's system.

It was argued that he didn't circumvent, but it didn't stop them from putting him in jail initially.


His conviction was vacated and Amazon has deep pockets and diffusion of internal liability. The illicit state drug charges did not help his case.


Ruled by whom? The International Court of the Northern District of California?


In the mid 20th century some people believed urban motorways were "progress" and wanted to build them everywhere, see for example:

https://en.wikipedia.org/wiki/Futurama_%28New_York_World%27s...


Thank you for sharing this.

This vision is absolutely horryfying, yet at same time incredibly interesting.


If you wanna look into another example search for the Abercrombie Plan in Edinburgh: it was a very ambitious urbanistic plan to "modernize" this city. For instance they proposed to demolish all of the historical Georgian and Victorian buildings in Princes Street and replace them with brutalist buildings and a motorway.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: