Hyperthreading/SMT is a significant boon for heavily threaded workloads. What makes that such a win while Bulldozer's implementation of "two integer units sharing a front-end, cache and FPU" is supposedly so bad? Because that description makes the 8 core Bulldozers sound exactly like a 4 core with SMT.
>Hyperthreading/SMT is a significant boon for heavily threaded workloads.
That's hugely debatable and depends on SW workloads and the SMT implementation + CPU pipeline design.
In SMT the execution engines, ALUs, FPUs, and caches are completely shared. When one thread stalls waiting for RAM, the second thread sneaks into the idle execution units. At best, SMT yields a ~10% to 20% throughput boost over a single thread.
>Because that description makes the 8 core Bulldozers sound exactly like a 4 core with SMT.
It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.
> At best, SMT yields a ~10% to 20% throughput boost over a single thread.
Exactly, which is a significant benefit for how marginal the costs are.
> It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.
Then surely it should be even better than 4 cores with SMT?
If you're going to argue that the problem with Bulldozer was it's weird semi-SMT solution, you need to explain how it would've been better without it (aka as a regular quad core). Because even if it just gets the 10-20% performance improvements from being a form of SMT it would be better to have it than to not. And if you have lots of integer unit-bound threads, it should be even better than that.
>Then surely it should be even better than 4 cores with SMT? If you're going to argue that the problem with Bulldozer was it's weird semi-SMT solution, you need to explain how it would've been better without it
I explained all the bottlenecks of the architecture in a comment above, that the issue was more than 4-core +SMT instead of true 8 cores. Please read it.
I did read it. It's not clear from it why you think having 2 integer units per core in a SMT-like configuration makes it worse. If you have <=4 threads it doesn't matter, just schedule the threads on different proper cores. If you have >4 memory/FPU/front-end heavy threads you should see the same benefit as SMT. If you have >4 integer arithmetic-bound cores, you should see a significant benefit beyond what SMT would give.
Now the very long pipeline and high memory latency are obviously significant issues with the architecture but those seem disconnected from the 4-core+SMT issue? I'm not questioning those issues at all, it's just not the part of your comment which interested me
EDIT: okay so in this comment: https://news.ycombinator.com/item?id=49809017, you explain that there's actually a fairly large part of a core that's duplicated, not just two integer units. If each "core" gets its own integer unit, register file and L1 cache, you're actually paying a ton of die space for it, unlike SMT which is "free". I can totally get how that can be a terrible trade-off for most workloads if it all ends up mostly starved due to front-end/FPU/memory throughput.
Well, 'core' gets weird when we talk about Dozer. And where everything else had problems making it work.
AFAIR, a bulldozer 'module' has what is exposed to a core as two CPUs, but, per everything above, is two integer cores, one shared FPU core, and depending on the version of the arch, possibly shared fetch/decode/other resources between all of that. Also AFAIR the decoder sucked as far as being able to feed both the integer cores, and the integer cores were more anemic compared to what was in, say, a K10H Phenom.
Sorry, I edited my comment while you were writing. I had missed that it's more than just "one core with two ALUs". The more silicon you dedicate to this almost-but-not-quite-SMT solution, the worse of a trade-off it becomes in situations which don't benefit from it, and it sounds like quite a lot of silicon was dedicated.
Your theory that "1 shared FPU per module should equal 1 shared FPU per SMT core" makes sense on paper, but Bulldozer lost to Intel’s Sandy Bridge 4C/8T in floating-point and memory-heavy workloads because Intel's individual FPU, cache hierarchy, and front-end pipelines were vastly wider and faster than Bulldozer's shared components.
Having the same count of units (4 FPUs on the chip) did not mean having the same throughput. It's a HW bottleneck, not something AMD could fix via the OS's kernel allocation and scheduling of resources to the CPU to be able match Intel.
In strictly integer 4-8 thread benchmarks, yeah, AMD was often tied to Intel's 4C+SMT.
Bulldozer’s design didn't lose because the concept of sharing an FPU between two threads is worse than SMT. It lost because:
Intel's FPU was natively twice as wide (256-bit vs. split 128-bit).
AMD's write-through L1 cache caused catastrophic write contention in L2.
AMD's L2 and L3 caches had double to triple the access latency of Intel's.
A single shared 4-wide decoder couldn't feed an FPU and two integer units simultaneously.
There's nothing kernel developers can do to "not permit" kernel-level anti-cheat, anyone can build an anti-cheat kernel module (as long as you honor the GPL licensing requirements naturally).
Absolutely 100%. Now I'm very biased against this "AI" stuff in general, but even fans of "AI" can't want this, right? I mean I know some people who are heavy "AI" users and they all either have a setup with Cursor/Codex/Zed/VSCode/whatever or their preferred local model, they'd never ever want to use "whatever Apple happens to ship with the OS". Same with the model Google bundled into their browser, growing Chrome from roughly 200MB to over 4GB. Who's the target audience here?
I actually like AI tools, but I don’t use Apple Intelligence. It’s literally the worst in class. It feels very late 90’s Microsoft. They are shoving it on us when we don’t want it and it’s worse than everything else out there. It’s the Windows Media Player of the AI industry.
Yet something like "Call <name of a person>" will still immediately call someone with a completely different name without even letting me cancel, right?
Just today: “give me directions to the ferry terminal”. I’m on an island with one terminal, I shouldn’t have to be specific.
“Here’s a list of three ferry terminals that are on different islands, none of which are the island you’re on.”
“Give directions to the Orcas Island ferry terminal.”
“Here’s directions to the Port of Orcas Island, which is on the side of the island opposite the ferry terminal.”
If I weren’t driving a camper van on narrow roads, I would have taken her up on the “other island” option. “How do you propose we get there, Siri? Maybe start by driving to the g-ddamned ferry terminal?!”
Oh, nice, they added a confirmation screen before you call! I'm not sure when they're planning to roll that feature out to Europe but once they do, maybe Siri will be occasionally slightly useful when I need to call someone while my hands are dirty or while I'm driving.
I was around in the 90's too, which is why I can confidently say it is that bad. What Apple is doing with AI today is exactly the same formula as some of the worst stuff Microsoft put out back then. I'm not talking about the big things. More like the bundled apps where they see someone else getting traction with something and then they go and bundle their own K-Mart version of it that no one asked for.
> Siri AI is dramatically better than old-school Siri. [...] it just works.
Yeah, after 30 seconds, assuming it doesn't fail outright. The only thing I use Siri for is to set timers on my watch when cooking. This used to take 2-3 seconds at most and even worked completely offline on the watch itself. That functionality is now effectively broken for me.
> It’s the Windows Media Player of the AI industry.
I personally found WMP great. I never felt any urge to install alternatives (and the few times I did end up using another media player, I always thought it was a worse experience than WMP).
No, there are a tonne of Apple users that aren't developers, that will probably love this. They'll currently be using paid subscriptions to OpenAI, but only via the web app, and will happily switch to using Apple Intelligence.
The pitch isn’t that it will replace claude or codex runing 284 agents vibeslopping for me. It’s that Siri is pretty damn good now, with safe access the data I kinda wish could be accessed safely by an assistant.
It has a ways to go on some of the productivity stuff, and third parties have work to do, but that’s fine. They built a better foundation than anyone else has for its purpose.
All I really want out of Siri is something I can give actual natural language instructions to about smart home stuff, calendars, timers, "app stuff", without it actively leaking my info to other companies, and in theory this new Siri should be able to do that.
Kinda slow, though. I'm hoping that's just early scaling problems.
Real people don't need that. Siri has offered a modest set of voice-integrated features for over a decade, and made almost zero impact. Same for Google Assistant, which sucked less but was still useless.
The average smartphone owner cannot be trusted to switch away from privacy-degrading technologies even if the alternative is good. Look at Facebook, TikTok, YouTube, Spotify; it's a Keynesian beauty contest, and ChatGPT has more mindshare than Siri.
The selling point is being able to say things like "remind me about the next early release from school the night before" and it'll dig up the info from your emails and then create the reminder.
And people really trust the bot to do that stuff for them? They risk letting down their kid by not being there to pick them up just to save half a minute of finding the right email? That sounds crazy to me, creating a reminder isn't exactly hard
Parents—in many places—already let themselves rely entirely on driving their kids to every activity, which puts them at much higher risk of getting maimed in an accident under a guise of efficiency and spurious claims of other danger imposed by living closer to people and the activities that their kids could get themselves to. Many people are lazy, fearful, and prioritize comfort and control over all else, therefore it doesn't surprise me at all to hear people would opt for the convenience whenever an opportunity strikes.
This reads like a comment from a non-parent, or a comment from the kind of parent who argues its okay for their kid to spend 16 hours with their eyes glued to a screen because "blippi is educational".
> a comment from the kind of parent who argues its okay for their kid to spend 16 hours with their eyes glued to a screen because "blippi is educational".
So, just to make it abundantly clear that this is all a bunch of unfounded generalization: the kid in question takes the bus, can get home and get inside independently, and the reason I'm asking this convenient machine for a reminder is just to avoid being surprised.
Asking my phone for a convenient reminder makes me lazy, fearful, and prioritizing comfort and control over all else. Never change, HN. Actually, you know what? Please change, HN. This shit is tiresome.
No, I wasn't commenting about your kid, I was commenting on my generalized observations about the tendency for people, including parents, to accept the possible cost of risk when there's the prospect of even minor amounts of increased convenience. A friend of mine who doesn't earn much took out a large amount of debt to get a car so he could drive just himself to and from work even though there's a train that would get him there faster and more safely because he has weird prejudices and has control issues, which are not uncommon among parents who move to the burbs to raise their kids and keep a watchful eye in an isolated environment. My comment was about why I wouldn't be surprised by any shortcuts people take even if there's some risk involved, because being alive is time intensive and many people intentionally put themselves in situations where they feel they need a stack of shortcuts at their disposal to get even extremely trivial tasks done. Could be smoking or not going to the gym or not having a strong social circle or whatever.
And you were doing this in response to someone using the voice assistant on their phone to make a reminder about their kid getting home from school early.
I don't take transit because I don't like microtransactions. I'd rather have one car payment and be able to go anywhere I want whenever I want than have to pay a few bucks (or more) every time I move. A train ticket to visit some relatives would be like $70 while the same trip only costs me $20 in EV fast charging. (Would be less if I had access to home charging)
As an example, it costs like three dollars to board a bus here. The maximum cost per day is two of those, but come on. If I needed even just to take a bus to work and back each weekday, that's already over a third of my car payment, and I lose carrying capacity, climate control, social distancing etc.
Also I have to walk to/from bus stops, which I can't easily manage due to my disability. That's why I keep an electric scooter in my vehicle nowadays, but I highly doubt a bus would allow me to bring one of those on board. So at that point just use the electric scooter without bothering with the bus, except when I did that, I got hit by a pickup and needed to go to the hospital and counseling, etc.
I prefer to just pay in bulk. I care about more than simply raw cost, but the little microtransactions always seem to be a worse experience.
I mean it seems like you have a valid reason, but the raw cost in terms of just the loan payment isn't a trivial difference. "already over a third of my car payment" means your car payment alone is 3x the total cost otherwise, which is... a lot. Once you factor in maintenance, insurance, depreciation (which has been lower in recent years), it seems like it would be quite a leap if you didn't have a disability that made public transit more difficult.
I think the difference between transit methods within a city that has built-in alternative options has been interesting to examine now that I have a motor vehicle myself (that I got for free, but costs a lot to actually drive). I went without one for 8 years before someone offered it to me, and I do value the flexibility to occasionally get out of town wherever I want, but the tires alone will cost $2000 when I need to replace them, and the gas is crazy, so it's perfect in the sense that it sits there taking up street parking until I really need to get someone that a bike, scooter, bus, train, or walking doesn't work as well for. I'm not encouraged to use it for anything that it's really not needed for, and if it died tomorrow I wouldn't replace it with anything.
Relying on only the bus, even if the transit system is well-funded, has never been a particularly ideal situation, but the advantage is that I get that time back instead of being singularly occupied by driving. If there's traffic, I'd wildly rather be on the bus, but the train is always better.
It shows where it found the information. In this case, it included a link to a text message from the school which had this info, and quoted it so I could see that it got it right, and it shows the reminder it created so I could see that it's correct. I don't really have to trust it, I can see that it's working.
Maybe you have 5 relevant calendar events but it lists 3 and calls it a day. That is from my experience far bigger type of problem with AI. Hallucination by omission
That's exactly what big tech executives and AI bros don't get: It would be useful to use AI for the grudgy but verifiable sub-steps of such a task, e.g. have it help with the email search, and have it suggest a reminder. Or even better, suggest creating a reminder when receiving/reading the email. But leave the user in control, and provide evidence and transparency. Instead, many companies put their bets on the long-shot, do-it-all assistent AI, but technology is just not there yet, or maybe never will even.
I could definitely imagine using a system where I could type in Spotlight "that message about early release from school", hit enter, and have the e-mail or SMS or Signal message or whatever show up in front of me. Boring, non-sexy "better search" style "AI" which is actually useful but doesn't appeal to investors.
Did you go into Siri and join the waitlist for Apple Intelligence?
If not, you’re still using spotlight so the behavior is 100% expected.
Yes, Apple could do a better job making it more obvious but then again Apple Intelligence is Beta so I can see why they require a positive action right now to enable it.
I had better success with earlier versions. For me Siri has regressed so much and had become so unreliable, that having to constantly repeat or rephrase or delete whatever nonsense it was doing was so frustrating, I stopped using it years ago.
Its quite the paradox when Apple is pushing Apple Intelligence so hard, yet their assistant that's been around for 15 years can't do even basic things now.
And this right here is why I think Apple should have renamed the new stuff to something else. Go into Settings/Siri, enable Apple Intelligence (you will probably have to join the waitlist), give it a day or two to re-index and then try.
I think you will be astonished at the difference in behavior. It's night and day.
I remember chatting to some people behind Siri (it was an acquisition), and there was a feeling it was better before Apple nerfed and neglected it. Which is very Apple.
I can imagine lots of things- mainly around making short summaries so the user can choose things more easily, or maybe as a smart autocomplete.
I can imagine a home automation app helping you build a workflow with an LLM. Or an RSS reader giving you a short summary of each article. A health app being able to pull recipes into a standard format. Maybe a fitness app being able to search your workout history if you give it good tools.
Idk I usually use my phone for reading HN on the toilet and listening to podcasts while cleaning, not much else. I don’t know what else people even do with a phone lol. But those are things I’ve done in the past that LLMs could be helpful for!
Is this based on user feedback for your app? Do they want your app to feature home automation or making short summaries? Since you were using your app as an example I mean
Not OP, but in my case they are features we already provide that we can make cheaper (for ourselves) by doing it on the user's device rather than paying OpenAI for it.
I would probably called pro AI these days considering how much I've worked on getting the people in my organisation to adapt and learn how to use it, and how much I use it professionally. I've not always been pro AI, I'm sure my previous comments have been quite the rollercoaster over the past year. These days I'm part of a process which will probably replace quite a lot of jobs with AI down the line. The morals of which is a different discussion.
I don't use AI personally though. So when I found out that apple intelligence is currently taking up 10gb on my macbook m1 with the tiniest configuration at the time, that's massive. I've turned everything off and denied it access to training on any sort of apps and it's still using up 10gb of my 128. Meaning that Mac OS now takes up 53.76% of my drive. The AI I don't use being 12.8 of that. If Apple sold devices with large drives as default that might be less of an issue, but they don't. It's a little ironic, but I originally switched from Linux to Apple because I want my technology to work out of the box, but now it looks like I'll soon be going back for that same reason.
I played around with local image generation models not long ago, and I downloaded them myself along with the rest of the software I needed. Then again, the average Apple user seems to want to put themselves on a leash(noose?) with Apple on the other end.
The AI in Siri is actually really good. I can ask it something like "what is Judy's favourite food?" and it will do a full search of the databases of all the default apps and find that info.
As for the other AI features such as autocomplete, memoji, notification summaries... I'm really not a big fan at all.
And poor Judy thought for just a moment that someone actually cared enough about her to remember her favourite food. Sorry Judy, youre just not that important I suppose. :|
I have ADHD and am not going to remember a lot about Judy. But I will write down the things that I know she likes in a note, and I'll occasionally ask Siri to list what Judy likes when I'd like to do something nice to her.
1. Lose the user, make him struggle to find items in the Start menu (Windows), struggle to type words on the keyboard (iOS), or struggle to manage find their subscription (Atlassian),
2. ???
3. Profit
If it's a dark pattern, I'd like to know what step 2 is. Who knows, maybe the goal is to overwhelm the user and it's a spy operation from the KGB.
I think you may be misunderstanding what the Apple LLMs or models are used for. It almost certainly isn’t something you used to write code for development. It’s to power all the intelligence features on the phone like whenever you ask Siri a question.
But there are plenty of successful projects which would probably have been taken down if it wasn't for clean room RE. I mean just look at the clean room IBM BIOS clones from "IBM compatibles" in the early days of the personal computer.
The background level of software copyright legal actions is significant enough. If plane attacks happened that much then it would give us solid evidence of TSA effectiveness even if they never caught anyone directly.
If I understand correctly, traditional IPv6 flow is:
* A host configures its own IP address via SLAAC
* The host sends a packet to its gateway with some destination address
* The gateway forwards the packet to the Internet
* Eventually, a response packet arrives to the gateway
* At this point, the gateway does neighbour discovery to try to figure out how to send the packet to the host
* The gateway might drop the packet or delay forwarding it until neighbour discovery completes
Why couldn't we change the flow to:
* A host configures its own IP address via SLAAC
* The host sends a packet to its gateway with some destination address
* The gateway forward the packet, and at the same time starts neighbour discovery because almost all computers which send outgoing packets will eventually receive some incoming packet
* When the response packet arrives, neighbour discovery is likely already done, or if not it got a good head start
Isn't this the obvious solution which wouldn't require changes to hosts or new protocols, just a small tweak to the router? Usually, when there's a seemingly obvious simple solution to a real problem and that solution hasn't been implemented by any of the clever people working in networking standards, there's a good reason and the solution isn't as simple as it seems. So what am I missing?
Im mildly confused as well, there is an even more immediate shortcut that I've certainly implemented before. in arp its not unusual to to just create a ip->mac binding from the source information in the ethernet header. where this potentially breaks down if we start looking at issues of trust. but its already the case in ND that we trust the endpoint to have executed the state machine to search for duplicates. so what's preventing us from doing the same thing here? maybe just layering concerns?
This was my immediate thought as well. Though I have not thought through any of the details, it did occur to me that the information needed would be in the "source" section of the header. I wonder if it is too much work to validate it somehow before using it?
I haven't ever done any programming at this layer of the stack, so I'm purely spitballing.
Technically, there's no broadcast in IPv6, so the host is supposed to join the local multicast group and do the neighbor discovery flow to find the "on link" address. And it's not guaranteed that the network is "symmetric".
Technically, this is also true for IPv4. You can have a proxy-ARP host impersonating the sender, but since it had never been fully specced, nobody cares about this scenario.
I’ve always thought that IPv6 has dramatically worse layering than IPv4. In IPv4 over Ethernet, there’s ARP, which layers over plain Ethernet, and IPv4 sits on top of the combination of ARP+Ethernet.
In the IPv6 world, neighbor discovery is IPv6, but only sort of, because the participants don’t necessarily have real addresses. So it’s a mess.
IPv6 link local layers over Ethernet the same way Arp does. Both contain a source/dest MAC which is used for forwarding, both contain the relevant neighbor info. If anything, keeping the protocol's self-discovery messages wrapped in the protocol itself is actually cleaner layering at the cost of complexity (the extra link local signalling addresses).
It really is not. There's a whole morass with possibly overlapping "on link" networks that nobody can implement correctly on the first try.
Then there's this whole pretend "it's not broadcast but multicast" song-and-dance with ND in IPv6. In IPv4/ARP the separation is clean, and no lower protocol details leak into the IP layer.
Link-local addresses were also meant to be used for LAN-only apps. Except that it quickly turned out that you can't actually use them reliably because some interfaces (like PPP tunnels) do not _have_ MACs.
Morass, mess, broadcast/multicast, etc aside (seems more like complaints of complexity than layering), IPv4+ARP is the textbook example of a layering violation. When you do want to violate, having the L2 info in the L3 packet is still cleaner than L3 info in L2. One is a protocol carrying its own glue in itself, the other is a protocol using different protocols (per L2) to discover the glue the same way it could have itself anyways. It's certainly convenient of course, but that doesn't make it cleaner layering. It also gives a consistent answer for different L2s e.g. cellular links because of this.
Sticking L2 into L3 means that L3 needs the ability to communicate with nodes with as-yet-unknown L2 addresses and that L3 nodes that don’t have an L3 address yet need to be able to transmit L3 packets. Both of these are quite messy, and APR completely avoids these problems.
(I am not, however, defending DHCPv4 - that has some of the same problem.)
ARP does not avoid this problem at all, it broadcasts until enough L2 information is exchanged to unicast (which usually happens to also be the point the L3 information is resolved).
This is the same broadcast-then-unicast process ND uses, except ND can also start as a multicast forward if MLD is supported (naturally falling back to broadcast on the switch if not).
ARP is a protocol that makes perfect sense even when spoken by hosts that only know their own MAC addresses and do not yet know their IPv4 addresses.
IPv6 ND is IPv6 except it has the weird edge case in that it is spoken between hosts that may not know their own IPv6 addresses. So you end up with delights like the “unspecified address) built into IPv6.
If I'm 192.168.129.10 and I want to resolve who 192.168.129.17 is, I make an ARP with the destination as ff:ff:ff:ff:ff:ff. This is a placeholder L2 destination which just means "everyone". I likely need to do something completely different when not on Ethernet (which is surprisingly common when you get beyond PCs on a LAN) and that may or may not involve ARP but we'll stick with ARP on Ethernet for now.
If I'm 2600::10 and I want to resolve who 2600::17 is, the IPv6 destination for the ND packet is set to FF02::1:FF00:17. This is a union of the multicast range with part of the destination address (so the request can almost always only go straight to the 2600::17 node rather than using a placeholder to blast to everyone). If Ethernet is in use, the L2 destination is derived and set to 33:33:FF:00:00:17 by and for the same reasoning. Different addresses will be derived e.g. for 2600::18
If I don't know my address yet (say, for DAD in this example), ARP actually uses a second made up address "0.0.0.0" for the source IP which just means unspecified. In ND, I do the same to be able to DAD my link local address by saying I'm :: (also all 0s) but at least the destination is still not ff:ff:ff:ff:ff:ff. As a bonus, since ND only uses the link local address as the source for ND, DAD for the link local address is the only time the source address can be unknown. DAD for any number of unicast addresses will always have the link local to put as the source, even if they are not in the same subnet in the L2.
> In ND, I do the same to be able to DAD my link local address by saying I'm :: (also all 0s) but at least the destination is still not ff:ff:ff:ff:ff:ff
There is no real difference. All Ethernet packets that have bit 7 set in the first octet are broadcast. A packet to 33:33:FF:00:00:17 will be broadcasted across the LAN.
In practice, ND will flood the network just like ARP unless switches are configured to snoop on higher-level protocols (proxy ARP/ND).
Have you ever asked yourself why there would be 140,737,488,355,328 (half of all) MAC addresses reserved for broadcast if it had no utilities over setting ff:ff:ff:ff:ff:ff? You're correct about fallback replication behavior matching that of broadcast (though that has to do with participating in or snooping IGMP/MLD rather than proxy ARP/ND), I'm just not sure you are considering any implications beyond a single aspect of that one scenario in the above.
That bit is the I/G (individual/group) bit, not the broadcast bit. In switches/routers participating in multicast (IGMP/MLD or snooping of), it is used as the hardware key for the L2 multicast replication lookup. In switches/routers not participating in multicast, unique multicast groups still allow a dedicated MAC entry hardware trap to send the packet to the CPU for processing (VRRP, LLDP, STP, NDP, LACP, and more). Because ARP uses ff:ff:ff:ff:ff:ff you either need to use an ACL on the protocol type in the ingress pipeline or trap all broadcasts to the CPU (both are inefficient in their own ways). The same is true of the host NICs, regardless what the network gear is doing, who can filter all ND requests not to their address(es) by have a match on the ND multicast MACs relevant to the device be processed and then a larger deny matcher for 33:33:FF:xx:xx:xx just drop all others. Also, an ff:ff:ff:ff:ff:ff destination can never be eligible for multicast lookup (even if the switch/router is participating in multicast), so even it's still needed even if a given node might treat it similar to ff:ff:ff:ff:ff:ff.
But yes, if you ignore all of those other things and are in a network without MLD support it'll all fall back to ARP forwarding behavior with just a less generic placeholder address filling the bits. One of the great failings of IPv6 - its approach can be as inefficient as IPv4 in pathological scenarios.
> Have you ever asked yourself why there would be 140,737,488,355,328 (half of all) MAC addresses reserved for broadcast if it had no utilities over setting ff:ff:ff:ff:ff:ff?
Mostly because of a historic accident.
> Because ARP uses ff:ff:ff:ff:ff:ff you either need to use an ACL on the protocol type in the ingress pipeline or trap all broadcasts to the CPU (both are inefficient in their own ways).
Since you're talking about switches, they can just snoop on ARP and avoid broadcasts entirely. Some switches do that. And the last time I checked, multicast on most (all?) modern switches is also implemented by punting packets to the CPU.
This is where assumption fails, in the original formulation it was even called the multicast bit (instead of the I/G bit) and broadcast was considered a special subset of the multicast use case. Quite the opposite of how you have framed things as an accident of having so many broadcast addresses. (pdf warning) https://archive.computerhistory.org/resources/text/DEC/ether... ironically, this is
> Since you're talking about switches, they can just snoop on ARP and avoid broadcasts entirely.
ARP broadcast suppression is definitely a thing but it requires more than just snoop, you still need some form of replication of the information to the other switches in the network and you need the actual suppression+generation functionality (ARP snoop alone just lets an L2 switch build an ARP table, it doesn't define what to do with it). In the best case this is itself done via multicast, in a middle case it's thrown into BGP or similar and distributed that way (if all of your nodes are routers), and in the worst case it falls back to broadcast across the network for anything not known on a local port.
ARP broadcast suppression is also harder than with the multicast address for the reason above. Snooping also does nothing for the NICs connected to "basic" L2 switches not doing ARP broadcast suppression while the multicast MAC still does (even when not actually forwarded via multicast).
IPv4 is an example of _correct_ layering. The hardware address is a detail that does not leak into upper layers. It's confined purely to the network layer.
In contrast, with IPv6 the whole 64/64 separation is a result of leaking the MAC address into upper protocols. Indeed, MAC was supposed to be a part of the publicly visible IPv6 addresses for hosts!
A network layering violation is when a protocol at one layer relies on its information being carried in protocols on other layers. It's not just when the addressing bits happen to match between layers, which would be done by the host locally without a separate L2 protocol anyways. Nor was what you're discussing a requirement of IPv6, it was an optional addressing scheme. Nor did it take on as a popular option. Nor does it do anything to explain why IPv4 leaking address resolution down instead of self containing it is supposed to be a correct example.
> That's exactly what's happening in IPv6. The host address leaks information about the underlying hardware into higher-level protocols
I think there is still confusion what "A network layering violation is when a protocol at one layer relies on its information being carried in protocols on other layers" means. As a practical examples:
"Reading a book has a main character 'John' in it and deciding to use that as your name in your speech" is not a layering violation for speech. At no point does anyone need to read to understand your name is John while speaking with you nor does anything break when you change your mind and decide to be called Andsynstd even though it has never been written in a book written in a book.
"You can find my name if you read that book over there" is a layering violation. They have to stop using speech with you, switch to reading the book at a completely different layer of communication, and then suddenly start calling you John in speech even though it was never communicated in speech. If they just say "what's your name" and you say "John" they don't need to get any information from outside the network layer, regardless if the bits in your response also contained your L2 address or not.
In your example, that you read your hardware as one option to come up with your address does not force anyone on the network to use a protocol other than IPv6 to learn your address and talk with you. The litmus test for this is "if you replace Ethernet with a different L2 which can't transport any protocol but L3 protocols on top of it, can you still resolve addresses?" If the answer is no then it's handled externally, if the external handling happens on L2 then it's a layering violation.
> It was a requirement initially.
Not at all. From section 2.4.1 of RFC 1884 in 1995, which introduced the concept of IPv6 addressing architecture you can continue reading past the paragraph mentioning the example of a link-local derived address to see it was never the only example:
Another unicast address format example is where a site or organization requires additional layers of internal hierarchy. In this example the subnet ID is divided into an area ID and a subnet ID. Its format is:
| s bits | n bits | m bits | 128-s-n-m bits |
+----------------------+---------+--------------+-----------------+
| subscriber prefix | area ID | subnet ID | interface ID |
+----------------------+---------+--------------+-----------------+
This technique can be continued to allow a site or organization to add additional layers of internal hierarchy. It may be desirable to use an interface ID smaller than a 48-bit IEEE 802 MAC address to allow more space for the additional layers of internal hierarchy. These could be interface IDs which are administratively created by the site or organization.
> WTF is "leaking down"? The higher protocol levels are supposed to use lower protocol levels.
Hopefully this is already explained in the part about what a layering violation actually is, but the problem is indeed not related to IPv4 riding on top of an L2. Oblivious transport of higher layers is the point of abstracted layers. The problem is ARP, an L2 protocol, is not oblivious to the information of the layers above it, such as L3 IP information, breaking the abstraction. IPv6 corrected this, the neighbor exchange information is always encapsulated in an L3 packet.
"Layering violation" has a pretty clear meaning in CS. It means that a layer needs information from an upper layer for the system to work, or if a lower-level layer internal details are not abstracted properly.
For example, NATs are a layering violation because a router, which is supposed to work on the level of individual packets, needs to understand the details of sessions established in higher protocols (TCP, SIP, FTP, ...) and mangle the packets accordingly.7
The other way around is IPv6. The details of SLAAC that are driven by 64-bit MACs of the Ethernet layer. They make it impossible to use masks larger than 64 bits. The largest installed base of devices (Android) does NOT support DHCP, which is the only non-manual way to configure such addresses.
> The problem is ARP, an L2 protocol
And? What is your point? ARP is not a layering violation, it operates at the correct layer and properly abstracts it. MAC addresses are an internal detail of its functionality, they don't leak into upper layers.
Beep boop :). No, at least not last I checked. I'm just a guy who's day job was developing a NOS which targets both ASICs and a custom software-based forwarding pipelines at one of the main enterprise network vendors. Nowadays I'm PLM for it but kinda miss getting to spend years working with every single bit of these kinds of protocols.
> "Layering violation" has a pretty clear meaning in CS. It means that a layer needs information from an upper layer for the system to work
Maybe you're used to layering in areas of CS outside of networking? E.g. page 476 of TCP IP Illustrated by Fall and Stevens gives an example in the opposite direction than what you just said:
The careful reader will note that this causes a so-called layering violation. That is, the UDP protocol (transport layer) is directly processing bits “owned” by IP (network layer).
That said, you're correct NAT is still also another network layering violation driven by IPv4's limitations.
> The details of SLAAC that are driven by 64-bit MACs of the Ethernet layer.
MACs of the Ethernet layer are 48 bits.
> They make it impossible to use masks larger than 64 bits. The largest installed base of devices (Android) does NOT support DHCP, which is the only non-manual way to configure such addresses.
Android does not use the MAC address derivation mode of SLAAC, it uses randomized addresses mode of SLAAC for privacy. There are several such standardized modes for SLAAC which are not based on the link layer identifier. This should follow because Android's most common IPv6 interface is the cellular radio which does not even have an Ethernet MAC to derive from.
> And? What is your point?
The part you cut off: is not oblivious to the information of the layers above it, such as L3 IP information, breaking the abstraction.
This point will only make or not make sense once we agree what a layering violation in networking is, until then there's not really sense trying to debate it.
> Maybe you're used to layering in areas of CS outside of networking? E.g. page 476 of TCP IP Illustrated by Fall and Stevens gives an example in the opposite direction than what you just said
Yes. The interaction of UDP and IP _is_ a layering violation, just like NATs or even VPNs. IP and ARP are not.
> MACs of the Ethernet layer are 48 bits.
Bluetooth MACs are 64-bit.
> Android does not use the MAC address derivation mode of SLAAC, it uses randomized addresses mode of SLAAC for privacy.
Yep. So it wastes 64 bits of the address space essentially for no reason. There is literally no advantage of IPv6 ND over stateless IPv4 autoconfiguration, except that stateless IPv4 autoconfiguration doesn't waste half of the address bits.
> The part you cut off: is not oblivious to the information of the layers above it, such as L3 IP information, breaking the abstraction.
You probably misunderstand what "abstraction" _is_, then.
> Yes. The interaction of UDP and IP _is_ a layering violation, just like NATs or even VPNs. IP and ARP are not.
Are you saying L4<->L3 mixing is a layering violation but L3<->L2 mixing is not or is there a more detailed reason you're trying to give?
VPNs are not usually considered a layering violation in networking. They do encapsulate lower layers but they place no expectation protocols in other layers rely on the encapsulated data or vice versa. Again, the key being whether or not there is cross reliance on data between layers in the protocols, not whether or not the bits exist in the packet.
> Bluetooth MACs are 64-bit.
Sure, but Bluetooth is not Ethernet and work on any form of Ethernet or IP over Bluetooth (or even Bluetooth as an IEEE standard) was not started until several years after the IPv6 standards we're discussing were already finalized.
> Yep. So it wastes 64 bits of the address space essentially for no reason. There is literally no advantage of IPv6 ND over stateless IPv4 autoconfiguration, except that stateless IPv4 autoconfiguration doesn't waste half of the address bits.
I'd be glad to explain some of the actual reasons why we keep chasing a 64/64 split if you'd care to know. It has nothing to do with an alternate history where embedding MACs was a requirement, it has to do with other reasons which are still relevant today.
Keep in mind it's very much supported by IPv6 (and even many networks out there) to use something other than /64s. We just keep choosing to do so and use assignment methods which require so because it makes sense for other reasons more important than how densely populated the host bit portion is in a given subnet.
> You probably misunderstand what "abstraction" _is_, then.
Always a possibility :), I hope you keep the same possibility open as well.
Huh? No, the MAC was never a part of the publicly-visible v6 address.
I know you're talking about SLAAC, but SLAAC is just a convenient way of picking a unique address. Changing the address wouldn't result in e.g. the packet being sent to a different MAC. Even sending packets to link-local addresses still does NDP, rather than parse the MAC out of the address.
The 64/64-bit split was in fact a result of (then planned) Bluetooth having 64 bit MACs. Moreover, the initial IPv6 RFCs did not have privacy extensions for SLAAC: https://www.rfc-editor.org/info/rfc2464/#section-4
> Even sending packets to link-local addresses still does NDP, rather than parse the MAC out of the address.
Broadly it's true that historically there was this idea for ethernet networks at least. It was always optional though. Even in that long obsolete rfc2464 it's described as the way to do SLAAC which was optional even in 1998.
This kind of thing doesn't normally count as violation of layering though. In protocol design its common to leverage identifiers from lower layers for addressing. For example many workings of the internet would be hard to imaging with the rule that you could not use IP addresses and ports in upper level protocols (like DNS, P2P protocols, etc)
The early RFCs were written more informally, so it's hard to say what was optional. However, the consensus was that SLAAC was supposed to be the main way to configure IPv6, along with fully manual configuration.
> In protocol design its common to leverage identifiers from lower layers for addressing.
Yes, that's why my email has the IP address of the mail server. And why my WhatsUp contains the IMEI of my phone.
Notice how a) it's doing NDP, and b) the link-local is fe80::506c:e9ff:fe08:9ba3 while the MAC is 00:23:6e:5b:b8:2b? The "506c:e9ff:fe08:9ba3" part of the address isn't being treated as a MAC address by the protocol -- it's just some opaque bytes.
Yes, those bytes can be picked by looking at a MAC address, but that's only one way to pick them and the protocol doesn't treat those bytes as having any particular significance, and in particular it never assumes they contain a MAC or tries to use them as an actual MAC, so it doesn't qualify as a layering violation.
Even weirder: when the router forwards the packet from the host to the Internet, it already sees both the source IPv6 and MAC address, so it could store them.
Maybe there are some weird situations where a host that just got its own IP address starts proxying for a third node that wants return packets to asymmetrically bypass the host?
This is what I don't understand. I'll be the first to stand up and say there's a lot about IPv6 I don't know, but why can't/doesn't the router learn how to talk to the host when the host sends that outbound packet?
I'm guessing it's one of those completely over-engineered bits about IPv6 that is that way just because they wanted to engineer in so much complexity almost for the sake of it
Generally IP does not make that assumption. If I remember right (it's too early), having to go through neighbor discovery combined with a switch doing some special processing on ARP/ND packets protects against identity hijacking in the LAN.
Wild guess: packet forward is implemented in hardware while arp/nd is software, with probably some things (think "hardware interrupt" or something alike) that allows hardware to "call" the software stack (for instance, when the link-layer addr is unknown)
So, to implement what you said, we need more than a simple router upgrade: we'd need to change the hardware, so that when a packet is forwarded from a source that's not in the mac table, the software can (asynchronously) perform an arp/nd lookup
There is probably a world of issue behind that behavior, but I do not know
Great guess, not sure why it's at the bottom of the responses so far :).
You can either have the hardware do additional lookups for every packet it processes or you can just follow the normal process for the very first packet from that IP. Or, exactly what this article is about for the best of both worlds.
Advanced ASICs usually go down a different path of offering the ability to validate the ND process (to prevent spoofing) rather than doing even more to trust whatever is sent. This is compatible with the way GRAND works, making it a win-win-win approach.
Apparently it's not good to require the router to start a new multicast address resolution on seeing unknown sources. It could be spelled out more but the text "This is particularly relevant for anycast and proxy addresses, where more than one node may be capable of responding" says that in some circumstances a lot of link local addresses might correspond to a single global address, and there could be a lot of traffic generated. So it's better to have the host opt-in to the prepopulating of the STALE entries.
"STALE allows the router to use the information it has already learned without requiring a new multicast address-resolution operation. The router can subsequently verify reachability using the normal Neighbour Discovery mechanisms."
edit: actually the RFC goes over this option as well and the reasoning there is slightly different than above (and maybe even the blog post) - see 8.9 at in https://datatracker.ietf.org/doc/rfc9131/ . Also notable that the RFC is from 2021.
My assumption is this isn't the way it works because it would be doing work up front, when it's not clear that the return will be necessary at all (think UDP). The response could be some time in the distant future or never, keeping the mapping in memory could be a problem (IPv6 design is 30 years old... and fast memory was even more expensive back then).
TBH Tahoe feels like an especially buggy .0 still. I'm updating as soon as my laptop shows it as available in the hopes that it'll feel closer to a .1 than what Tahoe is.
This is what made me move away from Ubuntu. Use LTS and encounter issues with outdated packages? "Well duh, you're supposed to upgrade to interim releases if you need remotely up to date software". Use interim releases and encounter bugs? "Well duh, it's an interim release. Of course it's a buggy mess, nobody uses those"
Every Fedora release is intended to be solid and they come out twice a year.
I moved to Fedora from Ubuntu about a year ago for my laptop. My main motivation was not be defaulted to snap packages. I have had a great experience i.e it gets out of the way and it doesn't fall apart when I update stuff. I was worried about SELinux but find Fedoras defaults just fine and intuitive.
I’d imagine they’re live updating the library paths in the binary headers, so anything shorter or equal to what they’re using is a simple rewrite, but longer is more complex.
The present is built mostly on layers of the long-past :)
It makes me think of when Arch merged /usr/bin and sbin with /bin and sbin. Having them split made sense in a ton of scenarios that used to be very common, but increasingly the split was vestigial for most users.
Yea, my understanding is that the split between / and /usr used to be more or less: the drive they used for / ran out of space, so they mounted another drive as /usr. As a consequence, / became where you put stuff that was essential during early boot, while /usr was where you put everything else.
But these days, "early boot" is handled by initramfs and we've all got drives large enough to not need the split anyway.
reply