Control and imitation learning · Solo project
Constrained Rocket Ascent
A 2D control experiment for reaching 100 km while respecting dynamic-pressure and acceleration constraints.
The mission
The simulator asks a controller to reach 100 km while keeping dynamic pressure below 70 kPa and acceleration below 8.5 g. The implementation includes the vehicle dynamics, Gymnasium environment, training stages, and held-out evaluation.
Feasibility before learning
PPO trained from scratch repeatedly settled below the target. A deterministic q-bucket controller reproduced the same shortfall, showing that the original 20 t propellant vehicle was energy-limited under the pressure constraint.
Raising propellant to 25 t produced a safe deterministic ascent above 130 km. That changed the work from reward tuning to learning a demonstrated, feasible trajectory.
Imitation and guarded fine-tuning
Behavior cloning learned the demonstrations but drifted in states outside the expert trajectories. DAgger collected those visited states, relabeled them with the scripted controller, and retrained the policy.
PPO returned only as a warm-started fine-tune. Fixed-seed evaluation rejected candidates when success, q compliance, or g compliance dropped below configured floors.
Results
The promoted policy reaches 100 km on 990 of 1,000 held-out episodes, with dynamic pressure and g-load inside their limits on more than 99% of runs. I also put it through a 35-case stress suite covering wind, sensor noise, actuator lag, propulsion loss, and mass and drag errors. A robustness-tuned version held 100% q/g compliance while averaging 87.9% success across the suite.
Limits
This is an educational 2D simulation, not flight software. The result does not cover 6-DOF motion, real sensors, engine transients, structural modes, or a flight atmosphere. Three compounded propulsion-loss cases remain outside the simulated vehicle envelope.