Machine LearningDev4 min reading time

FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution

InfoQ (AI, ML & Data)
Read full post
UC Berkeley and MIT researchers developed FreeToken, an open-source inference engine enabling efficient Mixture-of-Experts (MoE) model execution on consumer hardware by dynamically co-scheduling CPU and GPU workloads to overcome PCIe bottlenecks.

More on this story


More in Machine Learning

Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

Covered by 2 sources
Machine Learning4 min read

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

The New Stack (AI)
Machine Learning4 min read

Salesforce introduces Enterprise AI Harness, AI Control Plane

SiliconANGLE