flat assembler
Message board for the users of flat assembler.

Index > Main > Is SIMD domain cross expensive? And OoO

Author
Thread Post new topic Reply to topic
ttd3v



Joined: 17 Sep 2026
Posts: 44
ttd3v 25 Sep 2026, 18:18
I remember reading that crossing domains (e.g, gpr to x87 related) was bad for performance, but is those like xword, yword, or zword registers via vpinsrt also expensive? I have made some macros to make a "stack" in some of those registers for parts of a codebase thinking that it'd be better than using push/pop as they'd use memory. Which is sound, but if the domain cross penalty is significant, the cost of moving to memory will be lower than using a few push/pop.

And is vpinsrt OoO friendly? Like, if I am inserting on a yword on different lanes, it should be OoO possible. But I am unsure, and I myself couldn't find sources for it. I searched through intel sources and couldn't manage to find data about it.

Idk if this behavior is documented and/or ensured, but just in case, I guess it is decent to ask.
Post 25 Sep 2026, 18:18
View user's profile Send private message Visit poster's website Reply with quote
revolution
When all else fails, read the source


Joined: 24 Aug 2004
Posts: 21155
Location: In your JS exploiting you and your system
revolution 25 Sep 2026, 18:51
All those metrics need to be measured because they are CPU dependant. CPU architectures can differ greatly from each other. Identify each target system and measure code in various paths to determine which works best for each system.

However IME push/pop are very efficient, even though they "use memory". The stack is the hottest place in the CPU and is right there in L1 cache (and on some CPUs can access the load/store buffers for even greater performance).
Post 25 Sep 2026, 18:51
View user's profile Send private message Visit poster's website Reply with quote
ttd3v



Joined: 17 Sep 2026
Posts: 44
ttd3v 25 Sep 2026, 18:55
revolution wrote:
All those metrics need to be measured because they are CPU dependant. CPU architectures can differ greatly from each other. Identify each target system and measure code in various paths to determine which works best for each system.

However IME push/pop are very efficient, even though they "use memory". The stack is the hottest place in the CPU and is right there in L1 cache (and on some CPUs can access the load/store buffers for even greater performance).


So the answer to whether that "stack in simd" is efficient would be: "uhh... depends"? Also, AI once mentioned a "stack engine", is it a canonical thing? Couldn't find data regarding such too. Perhaps that's just a rumor or practical thing that exist on some CPUs Razz
Post 25 Sep 2026, 18:55
View user's profile Send private message Visit poster's website Reply with quote
revolution
When all else fails, read the source


Joined: 24 Aug 2004
Posts: 21155
Location: In your JS exploiting you and your system
revolution 25 Sep 2026, 19:00
ttd3v wrote:
[So the answer to whether that "stack in simd" is efficient would be: "uhh... depends"?
Yup. It depends. There is no fixed answer that applies everywhere. There are only guesses and heuristics. The only way to know for sure if it is right for each system is to measure it and compare each option.
Post 25 Sep 2026, 19:00
View user's profile Send private message Visit poster's website Reply with quote
Display posts from previous:
Post new topic Reply to topic

Jump to:  


< Last Thread | Next Thread >
Forum Rules:
You cannot post new topics in this forum
You cannot reply to topics in this forum
You cannot edit your posts in this forum
You cannot delete your posts in this forum
You cannot vote in polls in this forum
You cannot attach files in this forum
You can download files in this forum


Copyright © 1999-2026, Tomasz Grysztar. Also on GitHub, YouTube.

Website powered by rwasa.