Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos Paper • 2308.09951 • Published Aug 19, 2023
Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation Paper • 2308.04549 • Published Aug 8, 2023
SongComposer: A Large Language Model for Lyric and Melody Composition in Song Generation Paper • 2402.17645 • Published Feb 27, 2024 • 1
Streaming Long Video Understanding with Large Language Models Paper • 2405.16009 • Published May 25, 2024 • 1
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Paper • 2407.03320 • Published Jul 3, 2024 • 94