Machine LearningDev9 min reading time

Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

MarkTechPost
Read full post
Researchers from UC Berkeley and UT Austin developed FreeToken, an edge-native MoE serving engine enabling large models like 753B GLM-5.2 to run on a single workstation GPU by elastically mapping computation across available hardware. FreeToken supports interactive speeds for models ranging from 35B on laptops to 753B on workstations, targeting solo developers and small teams needing cost-effective local inference.

More on this story


More in Machine Learning

Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

Covered by 2 sources
Machine Learning4 min read

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

The New Stack (AI)
Machine Learning4 min read

Salesforce introduces Enterprise AI Harness, AI Control Plane

SiliconANGLE