Output-Feedback Control of Linear Continuous-Time Systems Using Discounted Inverse Reinforcement Learning

Han Wu et al.

IEEE Transactions on Cybernetics2026https://doi.org/10.1109/tcyb.2026.3651519article
AJG 3
Weight
0.37

What the paper says

This article proposes a novel discounted inverse reinforcement learning (DIRL) algorithm for linear quadratic (LQ) control of unknown continuous-time (CT) systems with partially observable states and an unknown discounted value function. Existing DIRL methods predominantly rely on full-state feedback, limiting their applicability to practical scenarios where only input-output data are available. To this end, a state reconstruction method is designed for the system controlled by an expert using the measured desired output. Based on this, a model-free output-feedback (OPFB) DIRL algorithm is presented to iteratively solve the unknown value function and the corresponding optimal OPFB control policy equivalent to the expert control policy. The convergence of the proposed algorithm and the nonuniqueness of solutions are rigorously analyzed. Finally, comprehensive simulations reveal the effectiveness of the proposed algorithm in recovering the expert control policy and its superior computational efficiency compared to state-of-the-art (SOTA) methods.

1 citation

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1109/tcyb.2026.3651519

Or copy a formatted citation

@article{han2026,
  title        = {{Output-Feedback Control of Linear Continuous-Time Systems Using Discounted Inverse Reinforcement Learning}},
  author       = {Han Wu et al.},
  journal      = {IEEE Transactions on Cybernetics},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1109/tcyb.2026.3651519},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Output-Feedback Control of Linear Continuous-Time Systems Using Discounted Inverse Reinforcement Learning

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.37

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.16 × 0.4 = 0.06
M · momentum0.53 × 0.15 = 0.08
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.