RTX 4060 running the 35B model, 39 Tokens per second? Berkeley and MIT have joined hands to open-source FreeToken
2026-09-06 01:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
The University of California, Berkeley, and MIT have jointly opened-source a project called FreeToken. This project aims to enable consumer RTX 4060 graphics cards to efficiently run large language models with 35B parameters by optimizing quantitative algorithms and memory management techniques, achieving an inference speed of 39 Tokens per second. FreeToken plans to lower the hardware requirements for large models on consumer devices and promote the widespread use of AI models; the related code is now available for developers to use.
Berkeley and MIT have joined forces to open-source the FreeToken project, aiming to achieve an inference speed of 39 Tokens per second when the RTX 4060 graphics card runs a 35B parameter model. By optimizing quantitative algorithms and memory management techniques, this project significantly lowers the hardware requirements for large language models. FreeToken plans to promote the widespread use of AI models on consumer devices, and the relevant code is now available for developers to use.