Rendered at 21:14:20 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ack_complete 19 hours ago [-]
AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.
The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.
brohee 5 hours ago [-]
"The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs."
Are those not based on standard ARM cores or ARM documentation on its cores not deep enough for your purpose?
ack_complete 3 hours ago [-]
Both. The specific documentation I'm referring to is their optimization guides, such as the Cortex-A72 optimization guide:
This has detailed information on latencies and throughput, which are important when optimizing SIMD code. But ARM doesn't publish optimization guides for all their cores.
On top of that, the cores are often modified in significant ways. Snapdragon CPUs, for instance, have used modified Cortex cores in the past and can have performance differences from the original core.
To be fair, Intel's been slacking a lot on this too, not even bothering to update their own optimization guide for their latest cores. But that's made up for the community mining this information in great detail on sites like uops.info, and also being a lot less x86 cores to deal with. On the ARM side, there's practically not much analogous other than the Apple M1 microarchitecture analysis.
eqvinox 12 hours ago [-]
> some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions.
Or LSE, adding a whole different set of atomic ops... which are way faster on some CPUs...
officialchicken 16 hours ago [-]
I really hope this is my last x86-64 CPU, Intel has become an incompetent steward. The "killer feature" for AVX-2(56) at the time of original release was basically lag/jitter-free video playback. IMO, 512 should have never been released for desktop CPUs and restricted to servers. One day I will to switch to a mainstream Neoverse dev box running linux. I also target Cortex-M in rust, so it's got a lot of the typical issues related to missing docs (e.g. bringup of non-heterogenous cores, meaning that M3/M4 still can't be used in a big.little chip)
Tuna-Fish 10 hours ago [-]
Why? AVX-512 is the best SIMD ISA anywhere right now. And the width is its least important feature.
nextaccountic 5 hours ago [-]
It's not available everywhere, which means you can't count on it if you don't know your deployment target
0x457 3 hours ago [-]
Neither is AVX or AVX2. AVX512 isn't avaiable on some of the latest intel CPUs is purely because intel can't say goodbye to Skylake for some reason.
ack_complete 2 hours ago [-]
AVX2 is much more prevalent than AVX-512 and thus tenable to require. RHEL is switching to x86-64-v3 baseline, for instance, which requires AVX2. AVX-512, on the other hand, has been moving backwards since Intel has not shipped any consumer-level CPUs with it for years now. The Steam Hardware Survey has AVX2 at 95.4% while AVX512F is still far behind at 23.9%.
nixon_why69 15 hours ago [-]
> IMO, 512 should have never been released for desktop CPUs and restricted to servers.
Why not? To save die space?
i80and 9 hours ago [-]
I think they meant it should have been deployed to all SKUs instead of the very silly ecosystem split Intel did for a long time.
tancop 15 hours ago [-]
> ... you can just assert that they’re all recent enough to at least have AVX2 that was introduced over 10 years ago, and have the program crash or misbehave if it ever runs on anything without AVX2
> However, if you are distributing the binaries for other people to run, that’s not really an option.
This all depends on what kind of software you're making. A lot of games set their requirements about 5 generations back, like FC 27 where the minimum is a Ryzen 1600. That lets them use AVX2 unconditionally and prevent complaints from users who tried to run it with a super old CPU.
Then you get whole Linux distros like CachyOS and Clear (RIP) that rebuild the world for each architecture level and have them as separate variants. I think it still counts as binaries for other people.
swiftcoder 9 hours ago [-]
The other good(?) news on this front is that Windows 11 only supports chips with at least AVX2, so if you are willing to cut off everyone on win10 and earlier, you can stick to that
ack_complete 6 hours ago [-]
Pretty sure it only actually currently needs SSE4, since they started compiling it with POPCNT for x86. The official system requirement is Intel 8th gen CPU, but Intel made Kaby Lake and Coffee Lake CPUs that lack AVX.
sharktheone 24 hours ago [-]
I am hoping for portable SIMD so much. But I still think that often a manually rolled SIMD will be faster.
Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.
dwattttt 22 hours ago [-]
I guarantee you with my lack of skill, my attempt at using portable simd will exist while using manual intrinsics I doubt I'd get there.
The question for me is whether portable simd will result in faster code than plain auto-vectorisation; for the simplest loops auto has me beat (the few times I've tried it), but I imagine as the complexity grows I'll be more likely to try do something that breaks auto-vectorisation, and it'll be more obvious to me when I do that in portable simd.
nextaccountic 3 hours ago [-]
nowadays it doesn't matter if you personally don't have this particular expertise, people still expect you to use an agent to write this code
astrange 4 hours ago [-]
Portable SIMD is a bad abstraction. The only two good choices are "something like ISPC" (basically "autoscalarization", similar to GPU kernels) and "writing in assembly" (ffmpeg does this for a reason).
Archit3ch 3 hours ago [-]
Note that ISPC does not handle symbolic math e.g. sparsity preserving linear solvers or equation simplification a la Mathematica. Manual ASM does. ;)
cbolton 9 hours ago [-]
The handwavy section on RISC-V is underwhelming. Yes the ecosystem is still nascent but something like the SpacemiT K3 is far from "abysmal" for vector operations and would make a nice test case with its two core types both supporting RVV 1.0 (one with vlen 256, the other with vlen 1024). On the more industrial side, for example the SiFive Intelligence X280 has been used in TPU by Google for years already and I doubt they are the only ones so calling the ISA completely irrelevant is a bit of a stretch.
Just for the sake of curiosity it would be nice to have a peek at what SIMD in Rust looks like on RISC-V today. Yes, even if it requires some "obscure compiler flags" for now (while we wait for the Oilsm extension).
camel-cdr 3 hours ago [-]
Also, rust could just be based and make -mno-strict-align the default.
9 hours ago [-]
aayush0325 8 hours ago [-]
I would love to see something like google highway for rust
(I am not connected to Fearless SIMD in any way other than as a user.)
the__alchemist 24 hours ago [-]
I am using my own lib, `lin_alg`, which apes core_simd for floating point values, and extends the concept to vectors and quaternions. I will eventually replace the floating point portions with core::simd upon its arrival in stable Rust.
Downside: It's currently x86 only.
imtringued 12 hours ago [-]
"Hardware that does arithmetic is cheap, so any CPU made this century has plenty of it. But you still only have one instruction decoding block and it is hard to get it to go fast, so the arithmetic hardware is vastly underutilized.
Warning, Nitpick. Saying "hardware [...] is cheap" and "instruction decoding is expensive" (implied) is a minor contradiction. You can actually duplicate instruction decoding just fine, it's the thing before it where things go to hell: Instruction fetching.
Nothing prevents you from building a computer that can fetch, decode and execute 16 instructions at once, assuming they don't all write to the same register.
But if you want to do the same add repeated 16 times you'll need 16 times more program memory and 16 wider read ports on your caches and so on. SRAM is really expensive so this strategy will waste a lot of area on memory that you probably didn't need in the first place. I say this as someone who had to design a chip in university and basically you couldn't even find the primitive CPU in-between the massive SRAM blocks. By reusing the same instruction you can now increase your compute to memory ratio in terms of area.
Just a heads up for people who want to know why SIMD is a thing. I'm not criticizing the article, I just want more people to realize the pain that SRAM represents to chip designers.
Archit3ch 1 days ago [-]
Hot take: there is no portable SIMD.
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
raphlinus 24 hours ago [-]
You've got a point but are overstating it considerably. There is a big gap between just autovectorization and the portable primitives a library like Highway or Fearless SIMD will give you. For example, I haven't seen autovectorization do select or swizzle.
But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.
So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.
Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.
As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.
Asmod4n 23 hours ago [-]
The issue with libraries which offer you portable simd is that the auto vectorizer of the compiler will likely generate faster code.
simonask 14 hours ago [-]
This isn't true, often even in trivial cases. Auto-vectorization is actually quite fragile in 2026. The reason is that it's subject to (a) scalar float semantics (i.e. the resulting code must not produce different results from the scalar version), and (b) a number of opaque compiler heuristics that sometimes work out, sometimes don't.
For example, consider you want to compute the average of a list of floats. The compiler cannot autovectorize this, because float addition is not commutative. However, it's much faster to do component-wise addition in groups, then a horizontal sum at the end, and then divide. Whether it matters depends on your use case, and the compiler unfortunately can't read your mind, so it has to be conservative.
astrange 4 hours ago [-]
If you try writing SIMD by hand autovectorization can mess it up, eg if you have to write a scalar trailing loop then it might try to autovectorize it.
vlovich123 23 hours ago [-]
The autovectorizer afaik rarely emits optimizations for the different vector units to support + efficiently caches the CPUid check to happen once on program start. It’s a good baseline but the continuum (today) is scalar -> auto vectorized -> portable SIMD -> hand rolled explicit. That portable SIMD lets you bridge into hand rolled explicit ergonomically is a power auto-vectorization doesn’t have. Either the compiler does it or doesn’t but you have no way to even have a check that says “fail to build the program if this function isn’t vectorized”. This is important if you’re relying on that property and someone accidentally adds a data dependency and breaks the optimization without you realizing. Portable and explicit SIMD don’t have this problem by definition.
Asmod4n 19 hours ago [-]
Well, let’s hope rust is better than std::simd from c++ here which lacks the capability to emit some intrinsics.
Gcc and llvm can tell you if they can’t Auto vec a function, maybe rust could turn this into an error at comptime.
throawayonthe 11 hours ago [-]
but like, this isn't the case
horseloverthin 22 hours ago [-]
[dead]
Sharlin 1 days ago [-]
Getting 2x or 4x performance in your inner loops using a reasonable SIMD library is infinitely better than theoretically getting 8x performance with hand-coded nonportable intrinsics, because the latter is never going to happen in most programs, so the actual point of comparison is scalar code, or autovectorized code at best.
Archit3ch 12 hours ago [-]
2x/4x/8x is still thinking in terms of autovec.
I'm getting 50x faster code with manual ASM. That's the difference between audio code that runs in realtime and code that does not.
The competing implementations use SIMD and native code, autovec works nicely there. I symbolically invert the LinAlg system at compile time.
Pannoniae 22 hours ago [-]
No it's not because it sucks the air out from the actual solution. ISPC more than a decade ago managed to demonstrate close-to-linear speedups for increasing vector sizes, even for branchy code.
Nowadays you can even get AI to write intristics and it works just fine, the portable libraries/autovec aren't really a serious player here.
Portability is also overstated - see the recent shift where Spotify decided to make native Android/iOS apps again instead of React Native. Usually, the number of relevant platforms is somewhere between 2 and 3, so portability concerns are more theoretical than real.
adgjlsfhk1 8 hours ago [-]
> Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.
kccqzy 3 hours ago [-]
Have you used the highway library? It’s portable SIMD done right. It does not rely on auto-vectorization. Instead it gives you a nicer API than using intrinsics, plus machinery to do dynamic dispatch.
the__alchemist 24 hours ago [-]
I agree with that historically auto-vertorization does not seem to work reliably. I'm not sure about your broad claim.
Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.
That still just gets you autovectorization, and generally locks you out of the performance you could have with direct SIMD intrinsics.
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
louthy 12 hours ago [-]
> That still just gets you autovectorization
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
Which JIT languages use much simd in generated code, in general?
louthy 13 hours ago [-]
C# definitely does [1] and has intrinsics that allow auto-fallback (in the JIT compilation phase) to the most supported SIMD instructions (and in-software implementation if non are); it has a lot of low-level coding primitives that are picked up by the JIT compiler and optimised. I would assume JVM based languages too.
Definitely, and in Java's case there are plenty JVM implementations to chose from, which widens what is possible.
louthy 12 hours ago [-]
How does the JVM spot that the implementations should be converted to direct SIMD instructions? On dotnet, MS have the [Intrinsic] method-attribute that allows them to mark certain methods that the JIT knows about and do a compile-time replacement. With the JVM being a bit more 'general', how is that achieved? Or, do you just mean there's more implementations of the JVM itself and because of that there's choice?
pjmlp 10 hours ago [-]
I mean you get Hotspot on regular OpenJDK, Amazon and Microsoft OpenJDK forks have their own JIT downstream changes, Falcon on Azul, OpenJ9, PTC, Aicas,...
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
louthy 28 minutes ago [-]
Got it. Thanks for the clarification!
12 hours ago [-]
Tanjreeve 1 days ago [-]
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
louthy 13 hours ago [-]
That assumes no per-platform optimisation, which most JIT compilers will do. I can only speak to the dotnet platform as that's what I am using the most at the moment, but its JIT will spot certain functions, like Vector512.LoadAligned and replace it with the the SIMD equivalent. So, it's not just IR op-code to CPU op-code, it's looking for patterns-of-code that can be made more efficient at JIT compile-time.
xboxnolifes 24 hours ago [-]
Sure, in the same way that there is no such thing as portable code at all. The result will be suboptimal, but it will still be better than not having it.
MiroslavPokorny 17 hours ago [-]
Said it before and will say it again if binaries were distributed using a bytecode then the host o/s would and should be able to produce optimal binary when loading into memory.
MiroslavPokorny 18 hours ago [-]
This is going to be a real problem in the future when x86 and ARM, SIMD moves ahead and becomes wider.
Scene_Cast2 1 days ago [-]
What about numpy, numba, and torch.compile?
izacus 1 days ago [-]
Those are manually optimized per arch, aren't they?
Scene_Cast2 24 hours ago [-]
The libraries, yes, but the code you write is portable (at least until you get into squeezing the last few percent and switch to Triton / Helion in case of GPU, and even those are decently portable).
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.
IshKebab 1 days ago [-]
It's a continuum. Some things basically all SIMD implementations support. Want to add 2 4xf32 vectors together? That's pretty easy to do portably.
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
jedisct1 4 hours ago [-]
As someone who had pretty much given up on Rust after being very disappointed with its SIMD support, and had stopped following developments in that area, I was genuinely surprised to see someone announce this today:
The pattern provides a safe abstraction for runtime CPU detection and SIMD dispatch. It's gnarly AF, but Fearless SIMD 1.0 came out a few days ago so people won't have to implement it themselves anymore.
I'd recommend grabbing that lib and giving SIMD-in-Rust another shot.
The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.
Are those not based on standard ARM cores or ARM documentation on its cores not deep enough for your purpose?
https://support.arm.com/documentation/uan0016/a/
This has detailed information on latencies and throughput, which are important when optimizing SIMD code. But ARM doesn't publish optimization guides for all their cores.
On top of that, the cores are often modified in significant ways. Snapdragon CPUs, for instance, have used modified Cortex cores in the past and can have performance differences from the original core.
To be fair, Intel's been slacking a lot on this too, not even bothering to update their own optimization guide for their latest cores. But that's made up for the community mining this information in great detail on sites like uops.info, and also being a lot less x86 cores to deal with. On the ARM side, there's practically not much analogous other than the Apple M1 microarchitecture analysis.
Or LSE, adding a whole different set of atomic ops... which are way faster on some CPUs...
Why not? To save die space?
> However, if you are distributing the binaries for other people to run, that’s not really an option.
This all depends on what kind of software you're making. A lot of games set their requirements about 5 generations back, like FC 27 where the minimum is a Ryzen 1600. That lets them use AVX2 unconditionally and prevent complaints from users who tried to run it with a super old CPU.
Then you get whole Linux distros like CachyOS and Clear (RIP) that rebuild the world for each architecture level and have them as separate variants. I think it still counts as binaries for other people.
Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.
The question for me is whether portable simd will result in faster code than plain auto-vectorisation; for the simplest loops auto has me beat (the few times I've tried it), but I imagine as the complexity grows I'll be more likely to try do something that breaks auto-vectorisation, and it'll be more obvious to me when I do that in portable simd.
Just for the sake of curiosity it would be nice to have a peek at what SIMD in Rust looks like on RISC-V today. Yes, even if it requires some "obscure compiler flags" for now (while we wait for the Oilsm extension).
It's the closest thing to Google's Highway.
(I am not connected to Fearless SIMD in any way other than as a user.)
Downside: It's currently x86 only.
Warning, Nitpick. Saying "hardware [...] is cheap" and "instruction decoding is expensive" (implied) is a minor contradiction. You can actually duplicate instruction decoding just fine, it's the thing before it where things go to hell: Instruction fetching.
Nothing prevents you from building a computer that can fetch, decode and execute 16 instructions at once, assuming they don't all write to the same register.
But if you want to do the same add repeated 16 times you'll need 16 times more program memory and 16 wider read ports on your caches and so on. SRAM is really expensive so this strategy will waste a lot of area on memory that you probably didn't need in the first place. I say this as someone who had to design a chip in university and basically you couldn't even find the primitive CPU in-between the massive SRAM blocks. By reusing the same instruction you can now increase your compute to memory ratio in terms of area.
Just a heads up for people who want to know why SIMD is a thing. I'm not criticizing the article, I just want more people to realize the pain that SRAM represents to chip designers.
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.
So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.
Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.
As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.
For example, consider you want to compute the average of a list of floats. The compiler cannot autovectorize this, because float addition is not commutative. However, it's much faster to do component-wise addition in groups, then a horizontal sum at the end, and then divide. Whether it matters depends on your use case, and the compiler unfortunately can't read your mind, so it has to be conservative.
Gcc and llvm can tell you if they can’t Auto vec a function, maybe rust could turn this into an error at comptime.
I'm getting 50x faster code with manual ASM. That's the difference between audio code that runs in realtime and code that does not.
The competing implementations use SIMD and native code, autovec works nicely there. I symbolically invert the LinAlg system at compile time.
Nowadays you can even get AI to write intristics and it works just fine, the portable libraries/autovec aren't really a serious player here.
Portability is also overstated - see the recent shift where Spotify decided to make native Android/iOS apps again instead of React Native. Usually, the number of relevant platforms is somewhere between 2 and 3, so portability concerns are more theoretical than real.
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.
Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.
Except in languages with a JIT compiler
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
[1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
[1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
https://val.markovic.io/articles/philbin-the-safest-and-fast...
Performance is competitive with Zig, without using the "unsafe" keyword or nightly Rust.
Nice.
NGL, SIMD in Rust has not been a great experience. Philbin uses the "CPU Feature Tokens" pattern that the Fearless SIMD team (which the root article's author is a part of) came up with: https://shnatsel.medium.com/safe-simd-in-rust-even-on-the-in...
The pattern provides a safe abstraction for runtime CPU detection and SIMD dispatch. It's gnarly AF, but Fearless SIMD 1.0 came out a few days ago so people won't have to implement it themselves anymore.
I'd recommend grabbing that lib and giving SIMD-in-Rust another shot.