Startups Pursue Next Generation of Large Language Models Beyond Transformers

Companies are exploring alternatives to transformer architecture that powers current AI models, seeking improvements in efficiency and capabilities.

According to MIT Technology Review, several startups are working to develop alternatives to the transformer architecture that has dominated large language models since Google researchers introduced it in their 2017 paper “Attention Is All You Need.” The publication’s What’s Next series examines these emerging efforts to move beyond the current generation of AI technology.

The transformer architecture has been the foundation for modern LLMs including GPT and other prominent models. However, the article suggests that companies are now pursuing different approaches as they seek the next significant advancement in AI language technology. MIT Technology Review frames this exploration as part of broader industry trends shaping the future of artificial intelligence development.

The piece appears in MIT Technology Review’s What’s Next series, which the publication describes as examining industries, trends, and technologies to provide early perspectives on future developments. The article looks at how startups specifically are positioning themselves to potentially pioneer new architectural approaches that could succeed or complement the transformer-based models currently in widespread use.