MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks by from on 2023-05-04 22:40 (#6BF67) Comments