AI Researcher · LLM Post-training

Zhaoyuan Xia 夏照源

M.S. Student in Artificial Intelligence at Peking University · Algorithm Intern at Baidu

I build language-model systems that can read, reason over, and act on long, information-dense contexts with better evidence grounding.

Portrait of Zhaoyuan Xia
Building reliable AI systems

Profile

About me

I am a master's student in Artificial Intelligence at the School of Software and Microelectronics, Peking University, where I study long-context understanding, model post-training, and evaluation for language and vision-language models.

I am currently an algorithm intern at Baidu, working on long-context post-training for foundation models. Previously, I built prompt optimization agents at Baidu and geo-temporal reasoning benchmarks for vision-language models at SenseTime.

I received my B.S. in Computer Science and Technology from Central South University in 2024. I enjoy turning research ideas into measurable, deployable systems.

PyTorch Transformers SFT / RLHF / DPO RAG Agents

What I work on

Research interests

Long-context understanding

Evidence retrieval, information compression, citation grounding, and multi-document reasoning.

LLM post-training

SFT, GRPO, process rewards, and length curriculum for more capable and controllable language models.

AI agents

Self-evolving prompt optimization systems that diagnose errors, test hypotheses, and reuse experience.

VLM evaluation

Benchmarks and analysis for geographic, temporal, and cross-view reasoning in vision-language models.

Selected work

Publications

2026

GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models

Zhaoyuan Xia, et al.

A benchmark for geographic and temporal reasoning across maps, videos, multiple camera views, and unobserved regions.

Where I have worked

Experience

Baidu · Wenxin Foundation Model

Algorithm Intern · Long-context post-training

May 2026 – Present
  • Proposed the H2S reading paradigm and made evidence localization, summarization, and answer generation explicit in one autoregressive trajectory.
  • Built a structured training corpus across 11 benchmark families for long-context reasoning.
  • Trained and evaluated 7B/14B models for evidence-grounded long-context understanding.

Baidu · Mobile Ecosystem Group

Algorithm Intern · Basic search

Dec 2025 – May 2026
  • Built PEAgent, a self-evolving prompt optimization system with error clustering, hypothesis testing, slot-level edits, rollback, and candidate selection.
  • Designed teacher–student collaboration between strong models for diagnosis and planning and efficient models for batch evaluation.
  • Improved authority, satisfaction, and relevance metrics by 10.5, 13.3, and 4.0 percentage points and helped launch the internal platform.

SenseTime · Smart City Group

Algorithm Research Intern · Multimodal AI

Apr 2025 – Nov 2025
  • Led the development of GTR-Bench for geo-temporal reasoning across maps, videos, and non-overlapping camera views.
  • Evaluated more than 10 mainstream VLMs and analyzed gaps in spatial-temporal utilization, temporal prediction, and cross-view alignment.
  • The work was released with an open-source codebase.

Selected projects

Projects

Open source · 2026

Highlight-Then-Summarize

Long-context post-training framework and dataset for compressing distributed evidence before answering.

LLMGRPO128K context
View project
Research

GTR-Bench

Benchmarking geographic and temporal reasoning in vision-language models across multi-view environments.

VLMBenchmarkVideo + map
View project
Industry research

PEAgent

A prompt optimization agent that turns model errors into testable hypotheses and iterative improvements.

AgentsPrompt optimizationEvaluation
Deployed internally

Beyond research

Honors

2023National Second Prize · Service Outsourcing Innovation CompetitionTeam leader
2023Peking University Challenge Cup · First PrizeSpecial contribution track
—Peking University “Lixing” Program · Outstanding IndividualUndergraduate honors

Get in touch

Let’s talk about
reliable AI systems.

I am interested in long-context reasoning, post-training, evaluation, and practical systems that make model behavior easier to inspect.