To address the model fragility issues in long-distance visual-language-action (VLA) operations, researchers proposed a neural symbolic framework that integrates learned VLA control with explicit task graphs and multimodal process memory. This framework utilizes task graph encoding for action dependencies, effective transitions, and branching conditions, and maintains active steps, completed actions, text context, and relevant visual evidence through memory. Additionally, the study introduced spatial and temporal guidance through gaze or saliency cues from human demonstrations, and directly labeled pseudo-gaze points in teleoperation videos from a robot perspective to isolate their impact on strategy learning. The resulting guidance information was used for the fine-tuning and reasoning stages of VLA. Experiments were conducted in two long-distance operation domains: workspace cleaning and surgical instrument handling, covering metrics such as correct object and destination selection, subtask completion, task progress, step order consistency, overall task success rate, and process or execution errors. This work established structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-distance VLA operations.