Here is a brief, structured plan for testing long-context prompt chunking, designed to validate both functional correctness and retrieval