CMDR and CMDR-Bench address a challenge in searching multimodal documents: finding relevant pages for questions whose answers depend on information spread across a document. The work also describes an embedding model designed for this task.
The motivation is that retrieving pages in isolation can miss context needed to answer a query about multiple parts of a document. The material presents this problem and an evaluation approach; it does not, by itself, establish that the technique suits every collection or use case.
To study the method, gather representative documents and write queries whose answers require different pages. Record which pages should be retrieved and compare the results with page-level search, checking whether the necessary context appears—not just whether text matches.
Also test documents with varied formats and structures, then examine errors: missing key pages, irrelevant results, or fragmented context. If you use AI to summarize the paper or assess internal data, follow your organization’s policy: submit only authorized material and limit personal and confidential data.