检索结果-内蒙古大学图书馆

您好，读者！请登录

内蒙古大学图书馆

首页
概况
党建
资源
服务
科研支持
- 论文收录引用证明
- 科技查新
知识产权
档案馆
帮助

咨询与建议

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

您的常用邮箱：*

您的手机号码：*

问题描述：

当前已输入0个字，您还可以输入200个字

全部搜索
期刊论文
图书
学位论文
标准
纸本馆藏
外文资源发现
数据库导航
超星发现

高级检索

时间限定

出版年份：

文献类型

图书期刊文献学位论文多媒体

馆藏选择

电子馆藏纸本馆藏

核心期刊

全部期刊 SCI 收录期刊 SSCI 收录期刊 EI 收录期刊 CSCD 收录期刊 CSSCI 收录期刊

语言

中文英文

文献类型

期刊文献图书学位论文标准纸本馆藏

帮助

文字说明：

T=题名（书名、题名），A=作者（责任者），K=主题词，P=出版物名称，PU=出版社名称，O=机构（作者单位、学位授予单位、专利申请人），L=中图分类号，C=学科分类号，U=全部字段，Y=年（出版发行年、学位年度、标准发布年）

检索规则说明：

AND代表“并且”；OR代表“或者”；NOT代表“不包含”；(注意必须大写,运算符两边需空一格)

检索范例：

范例一：(K=图书馆学 OR K=情报学) AND A=范并思 AND Y=1982-2016
范例二：P=计算机应用与软件 AND (U=C++ OR U=Basic) NOT K=Visual AND Y=2011-2016

分类表

所选分类

>> <<

限定检索结果

文献类型

299 篇 会议
8 篇 期刊文献

馆藏范围

307 篇 电子文献
0 种 纸本馆藏

日期分布

学科分类号

180 篇 工学
- 158 篇 计算机科学与技术...
- 56 篇 电气工程
- 48 篇 软件工程
- 47 篇 控制科学与工程
- 13 篇 信息与通信工程
- 10 篇 机械工程
- 6 篇 仪器科学与技术
- 4 篇 力学（可授工学、理...
- 4 篇 生物工程
- 3 篇 动力工程及工程热...
- 2 篇 交通运输工程
- 2 篇 核科学与技术
- 2 篇 生物医学工程（可授...
- 1 篇 建筑学
- 1 篇 化学工程与技术
- 1 篇 航空宇航科学与技...
- 1 篇 食品科学与工程（可...
40 篇 理学
- 35 篇 数学
- 9 篇 系统科学
- 8 篇 统计学（可授理学、...
- 4 篇 物理学
- 4 篇 生物学
- 1 篇 化学
- 1 篇 天文学
- 1 篇 大气科学
- 1 篇 地球物理学
- 1 篇 地质学
18 篇 管理学
- 17 篇 管理科学与工程(可...
- 7 篇 工商管理
4 篇 经济学
- 4 篇 应用经济学
1 篇 医学

主题

115 篇 dynamic programm...
76 篇 reinforcement le...
67 篇 learning
47 篇 optimal control
30 篇 neural networks
27 篇 control systems
21 篇 approximate dyna...
21 篇 approximation al...
20 篇 function approxi...
20 篇 equations
17 篇 convergence
16 篇 adaptive dynamic...
16 篇 state-space meth...
16 篇 heuristic algori...
14 篇 mathematical mod...
13 篇 stochastic proce...
12 篇 learning (artifi...
12 篇 adaptive control
12 篇 cost function
11 篇 algorithm design...

机构

5 篇 arizona state un...
4 篇 department of el...
4 篇 school of inform...
4 篇 department of in...
4 篇 univ sci & techn...
4 篇 chinese acad sci...
4 篇 department of el...
3 篇 princeton univ d...
3 篇 northeastern uni...
3 篇 national science...
3 篇 robotics institu...
3 篇 univ illinois de...
3 篇 univ utrecht dep...
2 篇 univ groningen i...
2 篇 sharif univ tech...
2 篇 univ texas autom...
2 篇 pengcheng labora...
2 篇 guangxi univ sch...
2 篇 chinese acad sci...
2 篇 cemagref lisc au...

作者

14 篇 liu derong
9 篇 wei qinglai
8 篇 si jennie
7 篇 xu xin
5 篇 derong liu
4 篇 lewis frank l.
4 篇 martin riedmille...
4 篇 huaguang zhang
4 篇 jennie si
4 篇 marco a. wiering
4 篇 xin xu
4 篇 zhang huaguang
4 篇 dongbin zhao
4 篇 lei yang
4 篇 powell warren b.
4 篇 riedmiller marti...
3 篇 hado van hasselt
3 篇 van hasselt hado
3 篇 jagannathan s.
3 篇 munos remi

语言

305 篇 英文
1 篇 其他
1 篇 中文

检索条件"任意字段=IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning"

共 307 条记录，以下是251-260 订阅

全选清除本页清除全部题录导出标记到"检索档案"

详细简洁

排序：

An Online Model-Free reinforcement learning Approach for 6-DOF Robot Manipulators

An Online Model-Free Reinforcement Learning Approach for 6-D...

引用

2023 ieee international symposium on Robotic and Sensors Environments, ROSE 2023

作者： Hosny, Zeyad Nassar, Abdullah Aboelyazeed, Ahmed Mohamed, Mahmoud Abouheaf, Mohammed Gueaieb, Wail University of Ottawa School of Electrical Engineering and Computer Science OttawaONK1N6N5 Canada Bowling Green State University Robotics Engineering Bowling GreenOH43402 United States

ISBN: (纸本)9798350308044

Controlling 6 Degrees-of-Freedom (DoF) robotic manipulators in an online, model-free manner poses significant challenges due to their complex coupling, non-linearities, and the need to account for unmodeled dynamics. This paper introduces a model-free adaptive approach for real-time control of a 6 DoF 'EPSON' robotic manipulator, without requiring any prior knowledge of the manipulator's dynamics. Initially, we lay out the framework for an optimal control solution. A performance index is introduced, leveraging error dynamics and correction control signals, offering the capability to incorporate high-order error dynamics without the need to explicitly derive error trajectories. The order of error dynamics is determined by the chosen number of error samples. We assume a kernel-based solution structure aligning with the performance index, resulting in a temporal difference equation. This equation can be optimized to formulate a model-free control strategy. Subsequently, a reinforcement learning approach is adopted to approximate the underlying strategy. Infeasible exact solutions are overcome by employing a value iteration mechanism to adapt the actor-critic structures within an adaptive critics framework. To validate the proposed approach, it is compared against a conventional proportional-integral controller. A Unified Robot Description Format file is generated to facilitate the import of the robotic manipulator into the MATLAB Simulink environment, enabling its control. Ultimately, the proposed method yields superior results in terms of the dynamic characteristics of the response, demonstrating its effectiveness over the conventional approach. © 2023 ieee.

关键词： Real time control

来源：评论

学校读者我要写书评

暂无评论

Higher-level application of Adaptive dynamic programming/reinforcement learning - a next phase for controls and system identification?

Higher-level application of Adaptive Dynamic Programming/Rei...

引用

ieee symposium on Adaptive dynamic programming and reinforcement learning, (ADPRL)

作者： George G. Lendaris Systems Science Graduate Program Portland State University Portland OR USA

In previous work it was shown that Adaptive-Critic-type approximate dynamic programming could be applied in a “higher-level” way to create autonomous agents capable of using experience to discern context and select optimal, context-dependent control policies. Early experiments with this approach were based on full a priori knowledge of the system being monitored. The experiments reported in this paper, using small neural networks representing families of mappings, were designed to explore what happens when knowledge of the system is less precise. Results of these experiments show that agents trained with this approach perform well when subject to even large amounts of noise or when employing (slightly) imperfect models. The results also suggest that aspects of this method of context discernment are consistent with our intuition about human learning. The insights gained from these explorations can be used to guide further efforts for developing this approach into a general methodology for solving arbitrary identification and control problems.

关键词： Context Artificial neural networks Context modeling Process control System identification Training Humans

来源：评论

学校读者我要写书评

暂无评论

approximate dynamic programming for stochastic systems with additive and multiplicative noise

Approximate dynamic programming for stochastic systems with ...

引用

ieee international symposium on Intelligent Control (ISIC)

作者： Yu Jiang Zhong-Ping Jiang Department of Electrical and Computer Engineering Polytechnic Institute of New York University Brooklyn NY USA College of Engineering Beijing University China

This paper studies the stochastic optimal control problem with additive and multiplicative noise via reinforcement learning (RL) and approximate/adaptive dynamic programming (ADP). Using Itô calculus, a policy iteration algorithm is derived in the presence of both additive and multiplicative noise. It is shown that the expectation of the approximated cost matrix is guaranteed to converge to the solution of certain algebraic Riccati equation that gives rise to the optimal cost value. Furthermore, the covariance of the approximated cost matrix can be reduced by increasing the length of time interval between two consecutive iterations. Finally, the efficiency of the proposed ADP methodology is illustrated in a numerical example.

关键词： Noise Additives Symmetric matrices Covariance matrix Steady-state Approximation algorithms Convergence

来源：评论

学校读者我要写书评

暂无评论

dynamic lead time promising

Dynamic lead time promising

引用

ieee symposium on Adaptive dynamic programming and reinforcement learning, (ADPRL)

作者： Matthew J. Reindorp Michael C. Fu Department of Industrial Engineering and Innovation Sciences Eindhovan University of Technology Netherlands Robert H. Smith School of Business and Institute of Systems Research University of Maryland USA

We consider a make-to-order business that serves customers in multiple priority classes. Orders from customers in higher classes bring greater revenue, but they expect shorter lead times than customers in lower classes. In making lead time promises, the firm must recognize preexisting order commitments, uncertainty over future demand from each class, and the possibility of supply chain disruptions. We model this scenario as a Markov decision problem and use reinforcement learning to determine the firm's lead time policy. In order to achieve tractability on large problems, we utilize a sequential decision-making approach that effectively allows us to eliminate one dimension from the state space of the system. Initial numerical results from the sequential dynamic approach suggest that the resulting policies more closely approximate optimal policies than static optimization approaches.

关键词： Q factor Markov processes Schedules Nickel learning Supply chains

来源：评论

学校读者我要写书评

暂无评论

Adaptive optimal control for nonlinear discrete-time systems

Adaptive optimal control for nonlinear discrete-time systems

引用

ieee symposium on Adaptive dynamic programming and reinforcement learning, (ADPRL)

作者： Chunbin Qin Huaguang Zhang Yanhong Luo School of Information Science and Engineering Northeastern University Shenyang China Basic Experiment Teaching Center Henan University Kaifeng China

This paper proposes an on-line near-optimal control scheme based on capabilities of neural networks (NNs), in function approximation, to attain the on-line solution of optimal control problem for nonlinear discrete-time systems. First, to solve the Hamilton-Jacobi-Bellman (HJB) equation forward-in-time appearing in the optimal control problem, two neural networks are used to approximate the cost function and to compute the optimal control policy, respectively. And then, according to the Bellman's optimality principle and the adaptive technology, the on-line weight updating laws for the critic network and action network are derived, respectively. Further, considering NNs approximative errors, the stability analysis of the closed-loop system is demonstrated by Lyapunov theory. At last, a numerical example is provided to demonstrate the effectiveness of the proposed method.

关键词： Artificial neural networks Equations Optimal control Mathematical model dynamic programming Approximation methods Discrete-time systems

来源：评论

学校读者我要写书评

暂无评论

Development of reinforcement learning Algorithm for 2-DOF Helicopter Model

Development of Reinforcement Learning Algorithm for 2-DOF He...

引用

ieee international symposium on Industrial Electronics (ISIE)

作者： Andrew Fandel Anthony Birge Suruz Miah Electrical and Computer Engineering Department Bradley University Peoria Illinois USA

This paper examines a reinforcement learning strategy for controlling a two degree-of-freedom (2-DOF) helicopter. The pitch and yaw angles are regulated to their corresponding reference angles by applying appropriate actuator commands (input voltages) to the main and tail rotors of a 2-DOF helicopter using the proposed reinforcement learning [herein called the approximate dynamic programming (ADP)] strategy. Furthermore, the proposed strategy has the ability to configure the 2-DOF helicopter to track time-varying reference angles. The proposed ADP technique is capable of dealing with coupling effects between the rigid body structure and propeller dynamics associated with the 2-DOF helicopter model considered in this work. A set of computer simulations is conducted to evaluate the performance of the proposed algorithm. The performance of the proposed algorithm is also compared to that of a conventional linear-quadratic regulator (LQR).

关键词： Helicopters learning (artificial intelligence) Neural networks Mathematical model Approximation algorithms Rotors Adaptation models

来源：评论

学校读者我要写书评

暂无评论

Optimal tracking control scheme for discrete-time nonlinear systems with approximation errors

Optimal tracking control scheme for discrete-time nonlinear ...

引用

10th international symposium on Neural Networks, ISNN 2013

作者： Wei, Qinglai Liu, Derong State Key Laboratory of Management and Control for Complex Systems Institute of Automation Chinese Academy of Sciences Beijing 100190 China

ISBN: (纸本)9783642390678

In this paper, we aim to solve an infinite-time optimal tracking control problem for a class of discrete-time nonlinear systems using iterative adaptive dynamic programming (ADP) algorithm. When the iterative tracking control law and the iterative performance index function in each iteration cannot be accurately obtained, a new convergence analysis method is developed to obtain the convergence conditions of the iterative ADP algorithm according to the properties of the finite approximation errors. If the convergence conditions are satisfied, it is shown that the iterative performance index functions converge to a finite neighborhood of the greatest lower bound of all performance index functions under some mild assumptions. Neural networks are used to approximate the performance index function and compute the optimal tracking control policy, respectively, for facilitating the implementation of the iterative ADP algorithm. Finally, a simulation example is given to illustrate the performance of the present method. © 2013 Springer-Verlag Berlin Heidelberg.

关键词： reinforcement learning

来源：评论

学校读者我要写书评

暂无评论

Development of reinforcement learning methods in control and decision making in the large scale dynamic game environments

Development of reinforcement learning methods in control and...

引用

ieee international Conference on Computer-Aided Design

作者： S. Orafa M.J. Yazdanpanah C. Lucas A. Rahimikian M. Nili Ahmadabadi Control and Intelligent Processing Center of Excellence Faculty of Electrical and Computer Engineering University of Tehran Tehran Iran

In this paper, an analytical comparison is done between dynamic programming and reinforcement learning methods in dynamic two-player games. The emphasis is on the large number of states and actions available for each player and different conflictive optimization objectives of these games that make them complicated in modeling and analysis. Optimization and decision making is done through quantifying a modified Q-learning algorithm. By this method, it is shown that the information processing in large scale-long stage games will take shorter times and will result in lower decision costs whereas dynamic programming methods cannot handle them across long time-horizons

关键词： learning Decision making Large-scale systems Game theory dynamic programming Intelligent control Control system synthesis Equations Optimal control State-space methods

来源：评论

学校读者我要写书评

暂无评论

Directed exploration of policy space using support vector classifiers

Directed exploration of policy space using support vector cl...

引用

ieee symposium on Adaptive dynamic programming and reinforcement learning, (ADPRL)

作者： Ioannis Rexakis Michail G. Lagoudakis Department of Electronic and Computer Engineering Technical University of Crete Crete Greece

Good policies in reinforcement learning problems typically exhibit significant structure. Several recent learning approaches based on the approximate policy iteration scheme suggest the use of classifiers for capturing this structure and representing policies compactly. Nevertheless, the space of possible policies, even under such structured representations, is huge and needs to be explored carefully to avoid computationally expensive simulations (rollouts) needed to probe the improved policy and obtain training samples at various points over the state space. Regarding rollouts as a scarce resource, we propose a method for directed exploration of policy space using support vector classifiers. We use a collection of binary support vector classifiers to represent policies, whereby each of these classifiers corresponds to a single action and captures the parts of the state space where this action dominates over the other actions. After an initial training phase with rollouts uniformly distributed over the entire state space, we use the support vectors of the classifiers to identify the critical parts of the state space with boundaries between different action choices in the represented policy. The policy is subsequently improved by probing the state space only at points around the support vectors that are distributed perpendicularly to the separating border. This directed focus on critical parts of the state space iteratively leads to the gradual refinement and improvement of the underlying policy and delivers excellent control policies in only a few iterations with a conservative use of rollouts. We demonstrate the proposed approach on three standard reinforcement learning domains: inverted pendulum, mountain car, and acrobot.

关键词： Support vector machines Training learning Space exploration Probes Training data Markov processes

来源：评论

学校读者我要写书评

暂无评论

approximate dynamic programming of continuous annealing process

Approximate dynamic programming of continuous annealing proc...

引用

ieee international Conference on Automation and Logistics

作者： Yingwei Zhang Chao Guo Xue Chen Yongdong Teng Key Laboratory of Integrated Automation of Process Industry Ministry of Education Northeastern University Shenyang Liaoning China

approximate dynamic programming method is a combination of neural networks, reinforcement learning, as well as the idea of dynamic programming. It is an online control method which bases on actual data rather than a precise mathematical model of the system. This method is suitable for the optimal control of nonlinear systems, and can avoid the problem of dimension disaster. It can effectively solve the non-linearity of the plant or the uncertainty problem caused by the uncertainty of the system modeling. So, it is suitable for processing the complex system and task of time-varying. The heating section of the continuous annealing furnace consumes a large number of energy, and the dynamic programming method has some limitation for solve the problems. We design the optimization controller for the heating section of the annealing furnace based on the approximate dynamic programming method. In this paper, it mainly gives the basic structure and algorithm of the action-dependent heuristic dynamic programming method (ADHDP), and designs the temperature optimization controller of the heating section in the continuous annealing furnace based on the ADHDP method. Simulation shows the temperature controller based on ADHDP has some theoretical and practical significance for the future practical application.

关键词： dynamic programming Annealing Heating Furnaces Temperature control Uncertainty Design optimization Neural networks learning Mathematical model

来源：评论

学校读者我要写书评

暂无评论

没有更多数据了...

全选清除本页清除全部题录导出标记到“检索档案”

共31页 << < 22 23 24 25 26 27 28 29 30 31 > >>

检索报告对象比较合并检索0

隐藏清空

合并搜索

回到顶部

执行限定条件

内容：

评分：

请选择保存的检索档案：

请选择收藏分类：

订阅名称：

通借通还

温馨提示：

图书名称：

借书校区：

取书校区：

手机号码：

邮箱地址：

一卡通帐号：

电话和邮箱必须正确填写，我们会与您联系确认。

联系人：

所在院系：

联系邮箱：

联系电话：

内蒙古自治区呼和浩特市赛罕区大学西街235号邮编: 010021

建议与咨询 留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

时间限定

文献类型

馆藏选择

核心期刊

语言

文献类型

帮助

文字说明：

检索规则说明：

检索范例：

分类表

所选分类

限定检索结果

文献类型

馆藏范围

日期分布

学科分类号

主题

机构

作者

语言

请选择保存的检索档案： 新增检索档案 确定 取消

请选择收藏分类： 新增自定义分类 确定 取消

通借通还

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

请选择保存的检索档案：

请选择收藏分类：