Enable javascript in your browser for better experience. Need to know to enable it? Go here.

Research for reliable AI in critical systems

Thoughtworks AI Research Labs conducts research into how AI models can be evaluated, understood and controlled for use in critical environments at scale.

Our research

We focus on rigorous evaluation, interpretability, robustness and model control: how AI behaves, where it fails and how it can be guided. We also explore model interoperability and AI decision-making.

Read the research. Use the code.

Browse the Labs’ projects to access the research behind each one, with links to papers, code and datasets where available. Access the methods, run the code and build on the results.

More research

February 27, 2026

Concept consistency score

CCS measures how CLIP attention heads align with concepts; high-CCS heads preserve performance but can amplify social bias.
June 04, 2025

The next frontiers in AI — according to industry leaders

May 06, 2025

Calculating uncertainty in generative AI

October 31, 2024

LLM benchmarks, evals and tests

October 16, 2023

Decoding LLM uncertainties for better predictability

September 08, 2023

A surprisingly effective way to estimate token importance in LLM prompts

September 02, 2021

Probabilistic machine learning and weak supervision

September 01, 2021

A gentle introduction to machine teaching

Partners and collaborations

Thoughtworks AI labs sit within a wider network of organizations spanning public AI research, semiconductor innovation, cloud platforms, open source and AI engineering.

These relationships strengthen the lab’s ability to contribute to the methods, tools and technical standards shaping reliable AI.

 

For partnerships and collaboration inquiries

email ai-labs@thoughtworks.com