Subscribe Us

Ai2 Claims Its New AI Model Surpasses One of DeepSeek’s Best

AI2 Claims Its New AI Model Surpasses One of DeepSeek’s Best

Step aside, DeepSeek—there’s a new AI champion in town, and it’s American.

On Thursday, Ai2, a Seattle-based nonprofit AI research institute, announced the release of its latest AI model, claiming it outperforms one of China’s leading AI systems, DeepSeek V3.

According to Ai2’s internal testing, the Tulu 3-405B model not only surpasses DeepSeek V3 in several AI benchmarks but also outperforms OpenAI’s GPT-4o in certain areas. Unlike GPT-4o and DeepSeek V3, however, Tulu 3-405B is fully open-source, meaning that all components required to replicate it are freely available and legally permitted.

An Ai2 spokesperson emphasized that Tulu 3-405B demonstrates the United States' ability to lead in global AI development, particularly in generative AI.

“This is a significant moment for the future of open AI, reinforcing the U.S. as a leader in competitive, open-source AI models,” the spokesperson stated.
“With this launch, Ai2 introduces a powerful, American-made alternative to DeepSeek’s models—not just a major step in AI development but also a statement that the U.S. can lead in AI innovation independent of major tech corporations.”

AI2 Claims Its New AI Model Surpasses One of DeepSeek’s Best

The Power Behind Tulu 3-405B

Tulu 3-405B is a massive model. Ai2 reports that it contains 405 billion parameters, requiring 256 GPUs to run training processes in parallel. The number of parameters in an AI model typically correlates with its problem-solving ability—larger models tend to perform better than smaller ones.

One of the key techniques behind Tulu 3-405B’s success is a training method known as Reinforcement Learning with Verifiable Rewards (RLVR). This technique helps the model learn from tasks with verifiable outcomes, such as solving math problems and following precise instructions.

Beating the Competition

Ai2 claims that on PopQA, a benchmark consisting of 14,000 knowledge-based questions sourced from Wikipedia, Tulu 3-405B outperformed not just DeepSeek V3 and GPT-4o, but also Meta’s LLaMA 3.1 405B model. It also set a new record in GSM8K, a benchmark designed to test AI’s ability to solve grade-school-level math problems.

Tulu 3-405B is now available for public testing through Ai2’s chatbot web app, and its training code is accessible on GitHub and Hugging Face’s AI Dev Platform.

So, if you’re eager to test the next-generation open-source AI, now’s the time—before the next benchmark-breaking flagship arrives.

Post a Comment

0 Comments