Machine LearningDev9 min reading time

Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

MarkTechPost
Read full post
Researchers from UC Berkeley and UT Austin developed FreeToken, an edge-native MoE serving engine enabling large models like 753B GLM-5.2 to run on a single workstation GPU by elastically mapping computation across available hardware. FreeToken supports interactive speeds for models ranging from 35B on laptops to 753B on workstations, targeting solo developers and small teams needing cost-effective local inference.

More on this story


More in Machine Learning

Machine Learning3 min read

Anthropic caught scientists using Claude to further biological weapon research

Covered by 2 sources
Machine Learning3 min read

China's star AI labs routed user requests to Claude at least 35 million times in the summer: Anthropic

Business Insider
Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

Covered by 2 sources