Machine learning has advanced considerably in recent years, with systems surpassing human abilities in various tasks. However, the main hurdle lies not just in training these models, but in implementing them optimally in practical scenarios. This is where AI inference comes into play, arising as a primary concern for researchers and tech leaders al