Vals, an AI startup established in 2024, announced it has raised $40 million in a Series A funding round. Andreessen Horowitz led the investment, following a seed round previously led by 8VC and Bloomberg Beta. This funding will support Vals' mission to improve AI model evaluation.
The company was formed to address shortcomings in existing AI benchmarking systems. Co-founder Rayan Krishnan noted that many legacy benchmarks are not designed for modern AI models, leading to situations where companies can optimize models specifically for these tests. Vals aims to provide more robust and relevant evaluations.
Vals differentiates its approach by not publicly disclosing its specific test materials, which prevents models from being trained directly against the benchmarks. Instead of focusing on abstract intelligence tests, Vals evaluates AI models on their ability to complete complex tasks. This method seeks to verify what models can actually accomplish, rather than just their general knowledge.
The development of new, more secure, and task-oriented AI benchmarking could influence how AI capabilities are validated and advertised across the industry. As AI integration expands, accurate and reliable evaluation methods become more critical for verifying model performance claims.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Vals, a startup founded in 2024, secured $40 million in Series A funding led by Andreessen Horowitz to address limitations in current AI benchmarking systems. The company aims to provide more relevant and secure evaluations for modern AI models, moving beyond abstract intelligence tests to assess complex task completion.