Our research
We focus on rigorous evaluation, interpretability, robustness and model control: how AI behaves, where it fails and how it can be guided. We also explore model interoperability and AI decision-making.
Read the research. Use the code.
Browse the Labs’ projects to access the research behind each one, with links to papers, code and datasets where available. Access the methods, run the code and build on the results.
Post-training NVIDIA Nemotron 3.5 Lightning for enterprise domains
Published: August 11, 2026
EAGLE-3 speculative decoding for NVIDIA Nemotron 3.5 Lightning
Published: August 11, 2026
Publications and resources
Browse the full set of Thoughtworks AI Research Labs materials, including research papers, technical explainers, open-source code and datasets.
Hi-SEMFLOW: Lie algebra–based semantic flow for span-level informal language identification in Hindi
Rethinking skeleton-based action recognition from an action-class prediction distribution perspective
Do tokenizers fail on informal Hindi expressions? Evidence from static, downstream and robustness analyses
TinySQL: A progressive text-to-SQL dataset for mechanistic interpretability research
A semantic parsing framework for end-to-end time normalization
Cultural awareness in vision-language models: A cross-country exploration
MultiFeature graph convolutional network for OCR verification
Retrieval augmented forecasting for generalization
A superpersuasive autonomous policy debating system
Pruning the paradox: How CLIP’s most informative heads enhance performance while amplifying bias
Transformer-based temporal information extraction and application: A review
Antislop: A comprehensive framework for identifying and eliminating repetitive patterns in language models
p-less Sampling: A robust hyperparameter-free approach for LLM decoding
Beyond I am sorry, I can’t: dissecting large language model refusal
Towards transparent AI grading: Entropy as a signal for human-AI disagreement
Investigating the robustness of retrieval-augmented generation at the query level
Lie Algebra based semantic flow for incompleteness detection in summarization
Beyond linear steering: Unified multi-attribute control for language models
Training-free mitigation of language reasoning degradation after multimodal instruction tuning
Learning from reasoning failures via synthetic data generation
Is your paper being reviewed by an LLM? Benchmarking AI text detection in peer review
Scaling knowledge graph construction through synthetic data generation and distillation
Teaching a model to stop writing like a model
EAGLE-3 speculative decoding for NVIDIA Nemotron 3.5 Lightning
Post-training NVIDIA Nemotron 3.5 Lightning for enterprise domains
Cultural counterfactuals
Curveball steering — Geometry-aware non-linear steering to control LLM behavior
Anti-slopping — An innovation for rectifying LLM writing clichés
Concept consistency score
p-less sampling: A robust LLM decoding strategy
Steering smarter
Evaluating LLM-generated summaries using the Lie algebra framework
Distribution-aware feature selection for SAEs
Towards transparent AI grading: Entropy as a signal for human-AI disagreement
The next frontiers in AI — according to industry leaders
Beyond linear steering: Unified multi-attribute control for language models
Calculating uncertainty in generative AI
TinySQL
Evaluating LLMs using semantic entropy
LLM benchmarks, evals and tests
Turning up the heat: Min-p samling for creative and coherent creative outputs
Decoding LLM uncertainties for better predictability
A surprisingly effective way to estimate token importance in LLM prompts
Probabilistic machine learning and weak supervision
A gentle introduction to machine teaching
Teaching a model to stop writing like a model
EAGLE-3 speculative decoding for NVIDIA Nemotron 3.5 Lightning
Post-training NVIDIA Nemotron 3.5 Lightning for enterprise domains
Cultural counterfactuals
Geometric curriculum coverage for detecting summary incompleteness via Hausdorff distance
Curveball steering — Geometry-aware non-linear steering to control LLM behavior
Anti-slopping — An innovation for rectifying LLM writing clichés
Hi-SEMFLOW: Lie algebra–based semantic flow for span-level informal language identification in Hindi
Rethinking skeleton-based action recognition from an action-class prediction distribution perspective
Do tokenizers fail on informal Hindi expressions? Evidence from static, downstream and robustness analyses
TinySQL: A progressive text-to-SQL dataset for mechanistic interpretability research
Concept consistency score
A semantic parsing framework for end-to-end time normalization
p-less sampling: A robust LLM decoding strategy
Cultural awareness in vision-language models: A cross-country exploration
MultiFeature graph convolutional network for OCR verification
Retrieval augmented forecasting for generalization
Steering smarter
A superpersuasive autonomous policy debating system
Pruning the paradox: How CLIP’s most informative heads enhance performance while amplifying bias
Transformer-based temporal information extraction and application: A review
Antislop: A comprehensive framework for identifying and eliminating repetitive patterns in language models
Evaluating LLM-generated summaries using the Lie algebra framework
p-less Sampling: A robust hyperparameter-free approach for LLM decoding
Beyond I am sorry, I can’t: dissecting large language model refusal
Distribution-aware feature selection for SAEs
Towards transparent AI grading: Entropy as a signal for human-AI disagreement
Towards transparent AI grading: Entropy as a signal for human-AI disagreement
Investigating the robustness of retrieval-augmented generation at the query level
Lie Algebra based semantic flow for incompleteness detection in summarization
The next frontiers in AI — according to industry leaders
Beyond linear steering: Unified multi-attribute control for language models
Beyond linear steering: Unified multi-attribute control for language models
Training-free mitigation of language reasoning degradation after multimodal instruction tuning
Calculating uncertainty in generative AI
Learning from reasoning failures via synthetic data generation
TinySQL
Evaluating LLMs using semantic entropy
Is your paper being reviewed by an LLM? Benchmarking AI text detection in peer review
LLM benchmarks, evals and tests
Scaling knowledge graph construction through synthetic data generation and distillation
Turning up the heat: Min-p samling for creative and coherent creative outputs
Decoding LLM uncertainties for better predictability
A surprisingly effective way to estimate token importance in LLM prompts
Probabilistic machine learning and weak supervision
A gentle introduction to machine teaching
Partners and collaborations
Thoughtworks AI labs sit within a wider network of organizations spanning public AI research, semiconductor innovation, cloud platforms, open source and AI engineering.
These relationships strengthen the lab’s ability to contribute to the methods, tools and technical standards shaping reliable AI.
For partnerships and collaboration inquiries
email ai-labs@thoughtworks.com