Open-Source Generative AI

Abacus.AI is committed to open-source AGI and has significantly contributed to open-source AI and LLMs. Our research is open-sourced, reviewed, and published in top AI and ML conferences.

Today our focus is fine-tuning frontier open-weight models into reliable AI agents. The refreshed Smaug line applies one methodology — human-curated, real-world agentic traces combined with synthetic data grounded in hard examples — to three open-weight bases: Smaug Flash on DeepSeek V4 Flash for always-on enterprise agents, Smaug Mini on Qwen3.8 27B for compact multimodal tasks, and Smaug Agentic on Kimi K3 at frontier scale. We release the strongest results back to the community on Hugging Face.

Our open-source contributions to LLMs have led to several other open-source labs adopting some of our techniques and pushing the boundaries of enterprise and SOTA AI.

Here are the key contributions from Abacus.AI to open-source:

Open Source AI

The Smaug Line

Smaug Flash

Smaug Flash is an agentic fine-tune of DeepSeek V4 Flash and the workhorse of the line. It is tuned for continuously running enterprise self-improving agents: loops that stay alive for hours, carry a long context from step to step, and spend most of their tool calls outside code — reading and writing documents, querying data systems, calling APIs, and driving automations. DeepSeek V4 Flash is fast, cheap and robust for frequently running agentic tasks but is susceptible to spins and confusion in long-context tool use; Smaug Flash was trained to make those loops faster and less prone to spins and stalls at maximum reasoning effort while keeping the cost and speed of the base.

Only the attention factor matrices are adapted — three LoRA adapters merged as full deltas — so the weights load exactly like the official release, with the same 1M-token context, and any serving stack that runs DeepSeek V4 Flash runs Smaug Flash unmodified. On our paired runs it gains +14.3 on LiveBench agentic coding, +13.7 on AutomationBench and +19 on NL2Repo-Bench over the base, and leads Claude Sonnet 5 on every benchmark we compared. Full results are on the research page and in the model card on Hugging Face.

Smaug Agentic

Smaug Agentic is the largest model in the line and the clearest demonstration of how the recipe scales: an agentic supervised fine-tune of Kimi K3, the frontier Mixture-of-Experts model from Moonshot AI, trained on filtered multi-turn, tool-using coding trajectories with reasoning tokens masked from the loss. Every architectural parameter is unchanged from the base, so any inference stack that serves Kimi K3 serves Smaug Agentic as a drop-in replacement, and interleaved thinking is preserved across turns.

On our own runs it improves on the published Kimi K3 figures on GPQA Diamond (94.1 vs 93.5), DeepSWE (69.9 vs 67.5), SciCode (60.8 vs 58.7) and LiveBench Agentic Coding (64.6 vs 62.2). The behavioral change is bigger than the deltas: runaway reasoning is suppressed while normal deliberation is preserved, and across 113 agentic coding tasks and more than seven hours of continuous work it ran a median of 78 agent steps per task with no infrastructure errors and no timeouts. Full per-benchmark notes are in the model card on Hugging Face.

Smaug Mini

Smaug Mini is an agentic fine-tune of Qwen3.8 27B, the compact member of the line. It targets multimodal use cases and smaller reasoning tasks — the one-off jobs that need to look at an image or a video, follow a precise instruction and call a few tools — and inherits the base's native image and video input. Applying the same methodology to a 27B dense model produces the same shape of result as on the larger bases: on our runs Smaug Mini leads Claude Sonnet 5, Claude Opus 4.6 and GPT-5.6 Luna on IFBench, AutomationBench and JobBench, at a size that fits on a single GPU.

Our Benchmarks

LiveBench AI

With LLMs training on web-scale data, test-set contamination is a pervasive concern in LLM evaluation that can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and they are not reliable for hard questions. For example, LLM judges make mistakes up to 40% of the time on challenging math and reasoning tasks.

To resolve this issue, we developed LiveBench, a contamination-limited benchmark that evaluates large language models on a variety of general intelligence capabilities including reasoning, coding, language understanding, data analysis, instruction following, and mathematics. By frequently releasing updated question sets and developing new tasks over time, we ensure that our results remain an accurate assessment of LLM capabilities as new models are released. Questions are constructed from a variety of recent sources such as research papers and news articles so they can be easily refreshed over time. LiveBench also uses only objective ground-truth judgment to ensure unbiased results.

Our leaderboard at livebench.ai presents results for all major model providers, both open source and proprietary. Since its release, LiveBench has remained one of the most popular LLM benchmarks and has been a featured result in major model release reports.
LiveBench Charts
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Schwartz - Ziv, Neel Jain, Khalid Saifullah, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum
Authors
ICLR Spotlight
Copyright © 2026 Abacus.AI. All Rights Reserved

Sign up to ChatLLM to proceed

Get more access to ChatLLM and unlock powerful AI Agent capabilities

Access to 100+ AI models including Fable 5.1, GPT 6 Astra and Seedream 2.0

Get Started
$10 $7
1st Month Discount First month then, $10/month
Models
100+ AI & Image Models
Vibe Code
Vibe Code Apps
General Purpose Agent
General Purpose Agent
CLI + CoWork
CLI + CoWork
SuperComputer
SuperComputer
Learn more