TEOChat: Large Language and Vision Assistant for Temporal Earth Observation Data
2024
TEOChat is a vision-language assistant designed for temporal Earth observation imagery. Built with a LLaVA-style architecture (temporally shared vision encoder + LLaMA 2 through an MLP projector), it supports instruction-following over image sequences rather than single frames. The project is trained with TEOChatlas, a 554,071-example temporal instruction dataset spanning dozens of EO tasks, and demonstrates strong performance on temporal reasoning and remote-sensing dialogue benchmarks.