← Back to projects

Control and imitation learning · Solo project

Constrained Rocket Ascent

A 2D control experiment for reaching 100 km while respecting dynamic-pressure and acceleration constraints.

Python · PyTorch · Stable-Baselines3 · Gymnasium · NumPy GitHub repo
99.0% success over 1,000 held-out episodes
100 km Kármán-line altitude target
99.4% · 99.7% q-limit and g-limit compliance in nominal evaluation
35 stress cases in the robustness suite

The mission

The simulator asks a controller to reach 100 km while keeping dynamic pressure below 70 kPa and acceleration below 8.5 g. The implementation includes the vehicle dynamics, Gymnasium environment, training stages, and held-out evaluation.

Feasibility before learning

PPO trained from scratch repeatedly settled below the target. A deterministic q-bucket controller reproduced the same shortfall, showing that the original 20 t propellant vehicle was energy-limited under the pressure constraint.

Raising propellant to 25 t produced a safe deterministic ascent above 130 km. That changed the work from reward tuning to learning a demonstrated, feasible trajectory.

Imitation and guarded fine-tuning

Behavior cloning learned the demonstrations but drifted in states outside the expert trajectories. DAgger collected those visited states, relabeled them with the scripted controller, and retrained the policy.

PPO returned only as a warm-started fine-tune. Fixed-seed evaluation rejected candidates when success, q compliance, or g compliance dropped below configured floors.

Results

The promoted policy reaches 100 km on 990 of 1,000 held-out episodes, with dynamic pressure and g-load inside their limits on more than 99% of runs. I also put it through a 35-case stress suite covering wind, sensor noise, actuator lag, propulsion loss, and mass and drag errors. A robustness-tuned version held 100% q/g compliance while averaging 87.9% success across the suite.

Limits

This is an educational 2D simulation, not flight software. The result does not cover 6-DOF motion, real sensors, engine transients, structural modes, or a flight atmosphere. Three compounded propulsion-loss cases remain outside the simulated vehicle envelope.