What exists
What I built
- Adaptive-agent environment with state, combat mechanics, and curriculum enemy pressure.
- Reward shaping around survival, combat, adaptation, anti-cowardice, efficiency, terminal, and opportunity signals.
- Gymnasium wrapper plus FastAPI/Gradio surfaces and optional Qwen/LoRA paths.
Architecture
- Mechanics update state and expose the next decision point.
- Curriculum enemy changes pressure over time.
- Reward components score behavior and expose reward-hacking risks.
- Training code stays outside the environment through the Gym wrapper.