Hacker Newsnew | past | comments | ask | show | jobs | submit | Lwrless's commentslogin

It's cool! Recently I built a space explorer-like WebGPU project inspired by Space Engine the game, and your implementation runs smooth with rich visuals, that's far superior to mine.

I do find it sometimes hard to navigate and has some visual glitches, but man it is good. I hope one day it becomes my go-to for space exploring, since Space Engine is Windows only and it does not run on my Mac. And building the space must be a huge amount of work, there's so much that has to go into it.


The "4x AI performance" smells suspicious to me, M5 Pro does have slightly higher memory bandwidth than the M4 Pro, but there are also the new Neural Accelerators in the GPU cores, I've always wondered what they actually do, since the SoC has AMX (Apple Matrix Coprocessor) since the original M1 chip era.

So is this performance gain probably due to the new Neural Accelerators?


Yes, the "neural accelerators" in the M5 GPU cores really do 4x pre-fill performance over M4. You can find benchmarks online since M5 has been out for months now. I assume the M6 has whatever the next-gen version of those is. These are not the "neural processors" that have been there since M1 (those still exist) but are additionl matrix math accelerators inside the actual GPU cores. Like "tensor cores" on Nvidia GPUs.


Thank you, found some benchmarks and they look really promising. To my understanding those Neural Accelerators are like AMX but for GPUs. With those accelerators the GPU performance on a M5 Max in LLM inferencing would totally be on par with a 5090, that's quite impressive!


Which inference benchmark are you looking at? The prefill speeds should be comparable on some LLMs, but the 5090 has much a higher theoretical max decode speed.


I found some on Reddit, and here are the two benchmark results from review sites: - against a desktop 5090: https://nanoreview.net/en/gpu-compare/geforce-rtx-5090-vs-ap... - against a laptop 5090: https://nanoreview.net/en/gpu-compare/geforce-rtx-5090-mobil...

With larger (also really fast) unified memory on the chip, it could easily load larger weights, and the thing is that the quantized results on the M5s are really good, I believe that for most local inferencing, users would be using INT4 (maybe more bits per weight sometimes) quantized models, that might be where the Neural Accelerators kick in. In raw power, a desktop 5090 easily outperforms the M5 Max, but for this specific use case, I believe that the M5 Max is good enough.


I have to say, this looks really well designed and easy to use, I want one! But why Asus? I really want to know how they landed on this.


> I really want to know how they landed on this.

RAM prices.


ASUS does make power banks, and the vast majority of the cost of any ebike is in the battery.


Those resets and the removal of the 5-hour usage limit are quietly anchoring me to a much higher usage baseline. I've stopped rationing and just spawn a bunch of agents to work at whatever pace I want, because there always seem to be more resets on the way (at least true for the last week). And now I am actually worried about that if one day they just stop doing so, my "normal" workflow will suddenly exceed the limit, and upgrading will feel like a step backwards.


All these resets are given because of rival competition. When they win the market, we are going to be squeezed $$$.


Can anyone explain how you “win” the market of super intelligence? Particularly with open weights models now rivaling the frontier, it seems like a race to the bottom even if the prices don’t yet reflect that.


"race to the bottom" is the negative framing of "competitive market prices".

So a company or union might say "this is a race to the bottom" when someone new enters their market, but to people buying their services this might be seen as welcome competition.

Do you actually see a negative impact from competition in this area? Or do you just mean competition will further reduce prices?


"race to the bottom" is also a response to a "market for lemons". It's not necessarily a good thing because while pricing drops to the floor so also does value to the customer in general. Usually it happens when price is very visible but the details of what the buyer actually receives are not.


You can only outlaw stuff in your own countrys..

And if cheaper access is an advantage, other countries will surpass you


Regulatory capture. You get them to outlaw the part of the competion (safety!) that is unwilling to pricefix and participate in your margin and market division agreements.


Don’t think of this as super intelligence. Think of it as vendor lock-in.

We have a good compare, which is cloud in the early 2000s.

Everyone (at the time) thought cloud compute costs would go to zero.

What most corporations didn’t realize is how entrenched your workflows and processes get when you adopt cloud and you become heavily locked in that ecosystem.

That same ecosystem lock-in is what the frontier labs are hoping for with AI.


Competitors run out of money and shut down, perhaps. Although I don't see how that happens for anyone but Google.


Do we all shift over to Chinese models.


Yes.


You patent and protect (as best you can) the missing ingredients needed to get to AGI. Only half joking and also scary to contemplate. Qualcomm’s CDMA patent is on example. ARM and Texas Instruments two other examples.


You hit the nail on the head. It's a race to the bottom.


Everyone is raising the bottom. Kimi got 60% more expensive during the 2.x cycle despite staying the exact same size.

Now K3 is almost 6x the cost of the original K2 checkpoint, and while the parameter count finally jumped, it's still an extremely sparse MoE and definitely does not cost 6x what the original K2 checkpoint did to host at scale.

Race to the bottom only takes real effect when there's a cap to the capabilities, otherwise everyone races to the bottom of a rising target (how economically valuable the tokens are)


But Kimi (at least say) will release their weights. So surely, if their prices are too high, somebody else will host it cheaper?


a) Why and b) With what compute?

"Why" as in, why take lower margins when Moonshot currently can't service all the demand for the model anyways. Based on past models no one is going to massively undercut Moonshot: few have the chops to serve it as efficiently as Moonshot and of those few, most of them don't go for being the cheapest, they go for being fast + reliable (think Together, Fireworks).

You get what you pay for applies very much with how many axes there are to serving these increasingly large models.

-

And for "with what compute": as the value of a token goes up, what people are willing to pay for compute is going up.

Every once in a while I'll see a story about falling rental rates, but with even slightly more established clouds I've been seeing availability get worse and worse over time.

I'm pretty sure the only reason the highly informal indexes don't reflect this is because every neocloud trying to cash in on an NVIDIA Inception discount kicks off by selling unrealistically cheap compute for a bit.


It's unlikely that a clear single winner will emerge in this competitive market.


Whats the bottom here tho? Its obvious what winning is.


only way you could win the market would be to own all the gpus


AGI will be insanely priced. LLMs are retarded childrin compared to proper intelligence. So this is just the entry level intelligence like thing. But I doubt there will be an AGI accessible by anyone.


If AGI comes to exist it won't be "priced" at all, since the lab that creates it will either quickly be seized by the gov't for national security, or they will become the most powerful organization in the world and have no need to sell services to other corporations. they will become the only corporation.


Good sci fi premise, but not at all how AGI will happen.

It’s not going to be a singularity at one moment of time. It’s not going to be instant runaway self-improvement, no matter what doomers and fetishists say.

It’s going to be gradual. We’ll see glimmers of AGI, and the “G” part will be about gradual broadening of domains and deepening of capabilities.

All of the coding harnesses are already using their own tools to self-improve, and the HITL component is getting less frequent and at higher levels of abstraction.

That’s how AGI gets here: very gradually, no hard takeoff, and nobody will be able to pinpoint when exactly it happened.

So: also no single lab with a massive advantage, no government takeovers. It’ll be a lot less dramatic than the extremes believe. IMO, of course.


The problem is... the moment someone gets it and it really solves hard problems and solves scifi level shenanigans, the given country that owns it, could gain such unfathomable lead above anyone else that it will almost surely lead to an all out war. I don't even know whether it is possible to conceal that you have such capability....

Imagine all the fear- and warmongering kingmakers and powerful individuals when they realize they have no power over anything or anyone.....so game over for them. They won't like it at all, at all.

Another problem, that in order to make it understand real life, it needs robots or humans wired into it (brain interfaces) in order to test certain things in the real world. And that is another level we know almost nothing about, at least on the surface.

PS: do these self improving harnesses even work?


Someone could have a small private breakthrough tomorrow that gives sample learning efficiency of the brain, online learning, and consolidated memories.


So, extend what they said to cover a 50 year timespan? "Instantaneous" wasn't mentioned or implied.


If it’s 50 years and gradual, there will be many places very close. The only “winner takes all” scenario is a sudden breakthrough when nobody else is close.

Because it’s a gradient, not a binary.


If gradually one of these labs ends up with the top model that makes all the others uncompetitive, what's the difference? Why do you see time as important if the end result is the same? Winner still takes all, no?


LLMs may be retarded but at least they know how to write "children"


That's all you could point out?


Perhaps.

But that hasn't happened, and it may or may not ever happen; we don't know the future. All we know is the past and the present.

And that today, we have tokens to burn.


If they try to engage in monopolistic pricing, I'll simply start paying for GLM-5.2, K3, DeepSeek et al.


I cannot recommend enough OpenCode Go, which for 10 bucks a month lets you use all of them with pretty hefty limits. Use it while it lasts


$5 with a referral code for the first month.

You get $60 worth of usage. It’s too good to be true, yet it is.


This is actually what I've started to feel a vague nagging concern about as well..

Say what you will, there's no way I'm going back to non-AI assisted coding. Even though I don't use AI to generate code or assets, it's great for reviews and brainstorming etc.

What if OpenAI/Anthropic decide to do a Netflix/Spotify move and pull the rug out from under us one day?

Like skyrocketing the price, or limiting peasants to older models (because Glorious Leader said so), or maybe some leak comes out that they've been spying on us all along.


It's a competitive market. Kimi K3 has almost reached parity with Opus.


This guy knows what he's talking about


Quietly eh?


Full disclosure: I work at Phi.

For context, Phi is an AI-powered Chromium browser for macOS. I think many people here may not be familiar with it. The closest reference point is probably Arc, although Phi's goal is slightly different: a browser that you and your agents can both use.

The main changes in version 2.0 are Spaces for organizing tabs and fully sandboxed Profiles, which keep cookies and sessions separate. In practice, that means things like multiple Google logins no longer collide. URL rules can also ensure that specific sites open in the correct space/profile.

The assistant now has whole-browser context, and the guardrail for removing sensitive data runs locally on your Mac instead of on a server. I recently worked on optimizing the local model path on the Apple Neural Engine because it seemed wasteful to leave that hardware unused.

I am also experimenting with a new memory visualization system (we call it Nebula). Rather than displaying browser memory as a list or force graph, it presents it as a spatial surface that you can explore over time. This feature is still evolving, but it has been interesting to work on.

Next, we want to enable agents such as Claude Code or Codex to control your actual browser window rather than a separate headless Chrome instance.

I'm happy to dig into any of this.


Got myself the $20 subscription and tried it out. The 5-hour limit runs out surprisingly fast. Quality is okay but it feels slow, and even with my $20 Claude subscription on Fable, the credit usage ends up being lower. Fable usually catches issues in my Opus 4.8-generated code that I'd miss otherwise, but Fugu didn't. Makes me wonder if it's really at the Fable level. Hard to see the value here.


I use my Flipper Zero weekly (or more frequent). This new model feels much more powerful than the ones based on RPi Zero as a handheld device. I like how they managed to include two RJ45 ports and a USB-A port for connectivity. However, it's still too bulky for me. Perhaps when I get one, I'll try carrying it around all day to see how it goes. There's also a nano SIM slot. With the two Ethernet ports, it's perfect for use as a mobile router. This use case alone is good enough for me.

For such a powerful device, I think the lack of a QWERTY keyboard and the inherited orange backlit monochrome display are two of its shortcomings. I don't want to carry a keyboard or screen with me, I want it to be able to take more human input/output without accessories.

For those interested in hackable, handheld Linux devices, the M5Stack Cardputer Zero is also worth a look. It will launch on Kickstarter soon, and I have reserved an early bird spot.


The keyboard on Cardputer is horrible, I mean you already doing a special edition, why not put a decent one (almost anything would be better)


I'm curious what this means for ChromiumOS and downstreams like FydeOS.

If Google is now pushing this "intelligence‑first" desktop experience, how much of that work is likely to stay in the proprietary ChromeOS/Googlebook layer vs. land in upstream ChromiumOS?


The OS on these Googlebooks will probably be a lot closer to Android 17 than to current ChromiumOS. Google has been consistent in saying that they're phasing out the ChromiumOS code base (while continuing the support the Chromebooks they've already sold) in favor of modifying AOSP to work better on laptops and desktops.


IMO this makes the argument for ChromiumOS and downstreams stronger. Gemini wants to do everything in my browser and i can't turn it off please help


I went a slightly different route. My switches are linked with 10Gbps SFP+ across the apartment, but it was way too late (and too much hassle) to pull proper in-wall fiber. Instead I used one of those ultra-thin (around 0.1mm), unshielded fiber cables, and just snaked it through door frames and taped it along the walls. I'm genuinely impressed this stuff exists, it makes retrofitting fiber into a finished space so much less painful.

Most of my edge devices are still on 2.5GbE though, and I'm increasingly aware that for anything with plain SATA disks, the drives are the real bottleneck. Once I LAG'd 2×2.5GbE to get a 5Gbps pipe, it became obvious the network wasn't the slow part anymore in a lot of cases.

And yeah, the 10GbE SFP+ modules run hot, so hot that I would not lay my fingers on them for more than 2 seconds. I stuck 2 copper heatsinks on my module, not sure they do much but the module runs smoothly. Even so, I'm pretty happy with the overall setup: from my 10GbE-equipped Mac I can saturate multiple machines at once and I no longer think about the network most of the time, which was the goal.


And after getting 10Gbps working at home I was getting greedy, and looked at InfiniBand as well, 40Gbps and proper RDMA is very tempting compared to Ethernet. The catch for me was the practical side: IB needs PCIe slots and those chunky, inflexible cables. With most of my stuff being laptops, mini PCs and Macs, I just couldn’t see a clean way to route those or even plug cards in everywhere, so in the end the "door-frame‑friendly" skinny SFP+ fiber still won out for this apartment.


What kind of fiber cables? If you have a link, I'd appreciate it. I can only find ones at 0.5mm.


Good catch — that was actually my mistake. I mixed up cm and mm from the marketing material, and when I actually measured just now it came out to around 0.9mm. So not quite as thin as I claimed, and I probably wouldn't have noticed if you hadn't asked.

I did find some Kevlar-reinforced options that are supposedly ~0.3mm, but they seem to be raw fiber without connectors, purposed for drones, and I'm not sure about global availability.


Ah, I see. Thanks for checking. I'm considering doing a similar run at home, so was curious if there were any options that thin. 0.9mm should be plenty to work with. :)


Yes, they are with one GPU core fused off. I came across a die shot on the internet[0], and the GPU cores look huge. With some rough calculation, I estimated that the GPU cores together take up about 15.5% of the die area. I don't know much about photolithography, but I assume the same percentage of single-defect dies could be limited to a single GPU core failure, it's actually pretty surprising that Apple can get enough of these "rare draw" chips to build and ship a real product.

If the shortage continues, I would expect that they start using fully functional A18 Pro chips with one GPU core disabled with software. It kind of reminds me of the AMD Athlon days when user could use a pencil to unlock extra cores.

[0]: https://chipwise.tech/our-portfolio/apple-a18-a18-pro-die-sh...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: