ScreenAI: Google Research’s 5B-parameter vision-language model for UIs and infographics
Google Research describes ScreenAI, a 5B-parameter vision-language model built on the PaLI architecture with pix2struct-style flexible patching, trained on a mixture of autogenerated and human-labeled screen and infographic data (including a novel Screen Annotation task). The team reports state-of-the-art results on UI/infographic tasks such as WebSRC and MoTIF and competitive performance on ChartQA, DocVQA and InfographicVQA, and is releasing three new evaluation datasets: Screen Annotation, ScreenQA Short, and Complex ScreenQA.