FiDO: Fusion-in-Decoder optimized for stronger performance and faster inference

Annual Meeting of the Association for Computational Linguistics (ACL), 2022

15 December 2022

Joshua Ainslie

Sumit Sanghai

Abstract

Fusion-in-Decoder (FiD) is a powerful retrieval-augmented language model that sets the state-of-the-art on many knowledge-intensive NLP tasks. However, FiD suffers from very expensive inference. We show that the majority of inference time results from memory bandwidth constraints in the decoder, and propose two simple changes to the FiD architecture to speed up inference by 7x. The faster decoder inference then allows for a much larger decoder. We denote FiD with the above modifications as FiDO, and show that it strongly improves performance over existing FiD models for a wide range of inference budgets. For example, FiDO-Large-XXL performs faster inference than FiD-Base and achieves better performance than FiD-Large.

View on arXiv

Comments on this paper