Machine learningProject record · Sep 2026

Machine Learning & Model Research

Regime-aware reinforcement learning

Recurrent PPO, portfolio environments, and related regime and fundamental-overlay studies. Each experiment retains its own evaluation boundary.

PyTorchGymnasiumPPORecurrentPPOJump modelsGMMStreamlitPythonTushareLogistic regressionFundamental factors
Why it sits here. Placed by the scope of implementation, the available contribution evidence, and the distinct technical capability it demonstrates.
03 / Somewhere between data and understanding.STUDY IN SPACE

01 / IMPLEMENTATION & CONTRIBUTION

What the work involves

Donald implemented the recurrent-policy portfolio framework and its starter. Related local TET studies add logistic overlays and fundamentals pipelines; the jump-model paper is a separate reference.

Technical depth

Recurrent policies, PID-Lagrangian drawdown constraints, holdings/turnover-aware simulation, online versus oracle regimes, event-gated entry, reporting-lag alignment, fundamental transforms and controlled signal ablations.

The project family

rl_trader_frameworklstm_ppo_regime_startertet-jumpmodels-tusharetet_fundamentals_tushare_a_sharejump_model_regime

02 / RESULTS

What came out of it

Portfolio environments combine PPO-PID drawdown constraints, turnover-aware simulation, regime modules, and diagnostics. The research explicitly separates causal online labels from smoothed oracle labels.

03 / SUPPORTING EVIDENCE

Follow the source

Implementation notes, project records, and supporting artifacts.

Source context & project scope

No production-readiness, profitability or live-computability certification. Reference-paper presence is not authorship or completed reproduction. Theoretical prototypes and deployed execution must remain separate.

Signal contract, holdings-aware environment, constrained PPO and regime-source distinction.

SOURCE · 2026-09-17

Concrete constrained RL agent module.

SOURCE · 2026-09-17

Concrete holdings and cost environment.

SOURCE · 2026-09-17

RecurrentPPO, expanding-window GMM and long-only-with-cash baseline.

SOURCE · 2026-09-17

Technical TET signals, jumpmodels and logistic overlay.

SOURCE · 2026-09-17

Availability-lag alignment and three-arm ablation design.

SOURCE · 2026-09-17

Quarterly/TTM and robust cross-sectional transforms implemented.

SOURCE · 2026-09-17

Reference artifact only; paper contents not independently audited in this pass.

SOURCE · 2026-09-17
CONTINUE IN MACHINE LEARNING & MODEL RESEARCH

X-Trend reproduction