Tech Meridian ← LIVE FEED
RU

NEWS · MODELS · #239

ScreenAI: Google Research’s 5B-parameter vision-language model for UIs and infographics

Google Research describes ScreenAI, a 5B-parameter vision-language model built on the PaLI architecture with pix2struct-style flexible patching, trained on a mixture of autogenerated and human-labeled screen and infographic data (including a novel Screen Annotation task). The team reports state-of-the-art results on UI/infographic tasks such as WebSRC and MoTIF and competitive performance on ChartQA, DocVQA and InfographicVQA, and is releasing three new evaluation datasets: Screen Annotation, ScreenQA Short, and Complex ScreenQA.

KEY POINTS

  1. Google Research describes ScreenAI, a 5B-parameter vision-language model built on the PaLI architecture with pix2struct-style flexible patching, trained on a mixture of autogenerated and human-labeled screen and infographic data (including a novel Screen Annotation task).
  2. The team reports state-of-the-art results on UI/infographic tasks such as WebSRC and MoTIF and competitive performance on ChartQA, DocVQA and InfographicVQA, and is releasing three new evaluation datasets: Screen Annotation, ScreenQA Short, and Complex ScreenQA.
  3. This matters because a compact multimodal model that better understands UI layouts and infographic visual language can enable scalable QA, navigation, summarization, and automatic dataset generation for human-machine interaction tasks.

WHY IT MATTERS

This matters because a compact multimodal model that better understands UI layouts and infographic visual language can enable scalable QA, navigation, summarization, and automatic dataset generation for human-machine interaction tasks.

SOURCES & TIMELINE

1