Research

Foundation models & publications

Multimodal foundation models for documents, images and interfaces, plus recent work on evaluating frontier language models.

Areas

Document understanding

DocFormer and DocFormerv2: transformers that fuse text, layout and pixels to understand documents end to end.

Vision language and agents

LaTr for scene text visual question answering, MixGen for multimodal data augmentation and compressed history states for faster web agents.

LLM evaluations

Benchmarks and graded transcripts that measure when language models exploit flaws in their reward signal or test cases.

Mentees & Collaborators

I am fortunate and proud to have worked with incredible researchers. I have learnt from them much more than what I could teach back.

Selected Publications

Preprint 2026 RewardHack-Bench: A Benchmark for Reward Hacking Detection in vLLMs with Price, Benton and Hubinger
ICLR 2026 ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
ACL Findings 2025 R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
AAAI 2024 DocFormerv2: Local Features for Document Understanding 100+ citations
CVPR 2024 Enhancing Vision-Language Pre-training with Rich Supervisions
WACV 2023 MixGen: A New Multi-Modal Data Augmentation 175+ citations
CVPR 2022 LaTr: Layout-Aware Transformer for Scene-Text VQA 140+ citations
ECCV 2022 YORO: Lightweight End to End Visual Grounding
ICCV 2021 DocFormer: End-to-End Transformer for Document Understanding 580+ citations
WACV 2021 Saliency Driven Perceptual Image Compression 70+ citations

Full list on Google Scholar.