Business

Deep Cogito Raises $43M for Self-Improving AI Models

San Francisco AI lab Deep Cogito has raised $43 million in Series A funding to scale its post-training engine, which helps artificial intelligence models learn and improve from their own reasoning.

Unite.AI23 hrs agoBusiness
Image: Unite.AI

Deep Cogito, a San Francisco-based artificial intelligence startup founded by former Google AI Search engineers Drishan Arora and Dhruv Malrana, has secured a $43 million Series A funding round. Led by TQ Ventures, the round included participation from Benchmark, Nexus Venture Partners, Atreides Management, South Park Commons, and cybersecurity firm Zscaler. This latest financing brings the company's total funding to more than $56 million, which will be used to expand its research, engineering, and computing infrastructure.

Rather than focusing on expensive pre-training runs, Deep Cogito is developing post-training methods like Iterated Distillation and Amplification (IDA) to help models improve their own capabilities. The startup has tested these techniques on its Cogito family of open-weight models, which range from 3 billion to 671 billion parameters, including 70B, 109B mixture-of-experts (MoE), 405B, and 671B MoE variants. With Cogito v2, the company reported that its 671B model produced reasoning chains roughly 60% shorter than DeepSeek R1 0528 while remaining competitive. Remarkably, Deep Cogito spent less than $3.5 million combined to train eight Cogito models across this entire parameter range. Its latest Cogito v2.1 671B model uses an open-licensed DeepSeek base model and applies process supervision to help the system identify productive reasoning paths.

For AI practitioners and enterprise developers, this approach offers a way to bypass the massive capital requirements of pre-training while still achieving highly capable, specialized models. Instead of relying on retrieval-augmented generation (RAG) to feed external data to a general-purpose model, businesses can use Deep Cogito's post-training engine to bake proprietary workflows, security data, and evaluation criteria directly into a model's weights. This allows enterprises to build, own, and run highly specialized models with lower latency and reduced token consumption, shifting the competitive landscape from raw compute scale to efficient, self-improving learning algorithms.

This is our own summary of reporting by Unite.AI

More in Business