NEWS · MODELS · #239
ScreenAI: Google Research’s 5B-parameter vision-language model for UIs and infographics
Google Research describes ScreenAI, a 5B-parameter vision-language model built on the PaLI architecture with pix2struct-style flexible patching, trained on a mixture of autogenerated and human-labeled screen and infographic data (including a novel Screen Annotation task). The team reports state-of-the-art results on UI/infographic tasks such as WebSRC and MoTIF and competitive performance on ChartQA, DocVQA and InfographicVQA, and is releasing three new evaluation datasets: Screen Annotation, ScreenQA Short, and Complex ScreenQA.
KEY POINTS
- Google Research describes ScreenAI, a 5B-parameter vision-language model built on the PaLI architecture with pix2struct-style flexible patching, trained on a mixture of autogenerated and human-labeled screen and infographic data (including a novel Screen Annotation task).
- The team reports state-of-the-art results on UI/infographic tasks such as WebSRC and MoTIF and competitive performance on ChartQA, DocVQA and InfographicVQA, and is releasing three new evaluation datasets: Screen Annotation, ScreenQA Short, and Complex ScreenQA.
- This matters because a compact multimodal model that better understands UI layouts and infographic visual language can enable scalable QA, navigation, summarization, and automatic dataset generation for human-machine interaction tasks.
WHY IT MATTERS
This matters because a compact multimodal model that better understands UI layouts and infographic visual language can enable scalable QA, navigation, summarization, and automatic dataset generation for human-machine interaction tasks.