检索结果-内蒙古大学图书馆

您好，读者！请登录

内蒙古大学图书馆

首页
概况
党建
资源
服务
科研支持
- 论文收录引用证明
- 科技查新
知识产权
档案馆
帮助

咨询与建议

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

您的常用邮箱：*

您的手机号码：*

问题描述：

当前已输入0个字，您还可以输入200个字

全部搜索
期刊论文
图书
学位论文
标准
纸本馆藏
外文资源发现
数据库导航
超星发现

高级检索

分类表

所选分类

>> <<

限定检索结果

标题

标题
作者
主题词
出版物名称
出版社
机构
学科分类号
摘要
ISBN
ISSN
基金资助
索书号

作者

作者
标题
主题词
出版物名称
出版社
机构
学科分类号
摘要
ISBN
ISSN
基金资助
索书号

文献类型

322 篇 会议
18 篇 期刊文献

馆藏范围

340 篇 电子文献
0 种 纸本馆藏

日期分布

学科分类号

288 篇 工学
- 248 篇 软件工程
- 232 篇 计算机科学与技术...
- 13 篇 电子科学与技术（可...
- 7 篇 信息与通信工程
- 5 篇 控制科学与工程
- 4 篇 机械工程
- 4 篇 生物工程
- 3 篇 生物医学工程（可授...
- 1 篇 力学（可授工学、理...
- 1 篇 动力工程及工程热...
- 1 篇 电气工程
- 1 篇 核科学与技术
- 1 篇 农业工程
- 1 篇 环境科学与工程（可...
53 篇 理学
- 49 篇 数学
- 4 篇 生物学
- 4 篇 系统科学
- 4 篇 统计学（可授理学、...
- 2 篇 化学
14 篇 管理学
- 10 篇 管理科学与工程(可...
- 8 篇 工商管理
- 4 篇 图书情报与档案管...
3 篇 经济学
- 3 篇 应用经济学
2 篇 法学
- 2 篇 社会学
1 篇 教育学
- 1 篇 教育学
1 篇 农学
- 1 篇 作物学

主题

54 篇 performance
48 篇 parallel process...
33 篇 algorithms
33 篇 parallel program...
27 篇 languages
25 篇 design
20 篇 parallel algorit...
20 篇 gpu
9 篇 experimentation
9 篇 measurement
7 篇 graphics process...
7 篇 theory
7 篇 parallel
6 篇 scalability
6 篇 mpi
6 篇 parallel computi...
6 篇 concurrency
5 篇 parallelism
5 篇 graph algorithms
5 篇 multicore

机构

7 篇 carnegie mellon ...
4 篇 indiana univ blo...
4 篇 shanghai jiao to...
3 篇 univ of tokyo
3 篇 tsinghua univ de...
3 篇 univ chinese aca...
3 篇 massachusetts in...
3 篇 univ illinois ur...
3 篇 swiss fed inst t...
3 篇 mit csail united...
3 篇 tsinghua univ pe...
3 篇 univ calif berke...
2 篇 ist austria klos...
2 篇 fudan univ sch c...
2 篇 georgetown univ ...
2 篇 univ wisconsin d...
2 篇 shanghai key lab...
2 篇 univ of wisconsi...
2 篇 tsinghua univers...
2 篇 shanghai jiao to...

作者

8 篇 blelloch guy e.
7 篇 chen haibo
6 篇 hoefler torsten
6 篇 garland michael
6 篇 zhai jidong
6 篇 shun julian
5 篇 sun yihan
4 篇 dhulipala laxman
4 篇 chen wenguang
4 篇 tsigas philippas
4 篇 tan guangming
4 篇 wang haojie
4 篇 nikolopoulos dim...
4 篇 mellor-crummey j...
4 篇 gu yan
4 篇 kennedy ken
3 篇 taura kenjiro
3 篇 li jiajia
3 篇 yonezawa akinori
3 篇 pingali keshav

语言

338 篇 英文
2 篇 其他

检索条件"任意字段=Proceedings of the 5th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming"

共 340 条记录，以下是121-130 订阅

全选清除本页清除全部题录导出标记到"检索档案"

详细简洁

排序：

相关度排序

相关度排序
时效性降序
时效性升序

Correct and efficient work-stealing for weak memory models 13

Correct and efficient work-stealing for weak memory models

引用

18th acm sigplan symposium on principles and practice of parallel programming, PPoPP 2013

作者： Lê, Nhat Minh Pop, Antoniu Cohen, Albert Zappa Nardelli, Francesco INRIA ENS Paris Paris France

ISBN: (纸本)9781450319225

Chase and Lev's concurrent deque is a key data structure in shared-memory parallel programming and plays an essential role in work-stealing schedulers. We provide the first correctness proof of an optimized implementation of Chase and Lev's deque on top of the POWER and ARM architectures: these provide very relaxed memory models, which we exploit to improve performance but considerably complicate the reasoning. We also study an optimized x86 and a portable C11 implementation, conducting systematic experiments to evaluate the impact of memory barrier optimizations. Our results demonstrate the benefits of hand tuning the deque code when running on top of relaxed memory models. © 2013 acm.

关键词： parallel programming

来源：评论

学校读者我要写书评

暂无评论

Lock-free channels for programming via communicating sequential processes 19

Lock-free channels for programming via communicating sequent...

引用

24th acm sigplan symposium on principles and practice of parallel programming, PPoPP 2019

作者： Koval, Nikita Alistarh, Dan Elizarov, Roman IST Austria Austria JetBrains Austria

ISBN: (纸本)9781450362252

Traditional concurrent programming involves manipulating shared mutable state. Alternatives to this programming style are communicating sequential processes (CSP) [1] and actor [2] models, which share data via explicit communication. Rendezvous channel is the common abstraction for communication between several processes, where senders and receivers perform a rendezvous handshake as a part of their protocol (senders wait for receivers and vice versa). Additionally to this, channels support the select expression. In this work, we present the first eficient lock-free channel algorithm, and compare it against Go [3] and Kotlin [4] baseline implementations. © 2019 Copyright held by the owner/author(s).

关键词： Locks (fasteners)

来源：评论

学校读者我要写书评

暂无评论

A parallel Sparse Tensor Benchmark Suite on CPUs and GPUs 20

A Parallel Sparse Tensor Benchmark Suite on CPUs and GPUs

引用

25th acm sigplan symposium on principles and practice of parallel programming (PPoPP)

作者： Li, Jiajia Lakshminarasimhan, Mahesh Wu, Xiaolong Li, Ang Olschanowsky, Catherine Barker, Kevin Pacific Northwest Natl Lab Richland WA 99352 USA Univ Utah Salt Lake City UT USA Purdue Univ W Lafayette IN 47907 USA Boise State Univ Boise ID 83725 USA

ISBN: (纸本)9781450368186

Tensor computations present significant performance challenges that impact a wide spectrum of applications. Efforts on improving the performance of tensor computations include exploring data layout, execution scheduling, and parallelism in common tensor kernels. this work presents a benchmark suite for arbitrary-order sparse tensor kernels using state-of-the-art tensor formats: coordinate (COO) and hierarchical coordinate (HiCOO). It demonstrates a set of reference tensor kernel implementations and some observations on Intel CPUs and NVIDIA GPUs. the full paper can be referred to at http://***/abs/2001.00660.

关键词： sparse tensors benchmark GPU roofline model

来源：评论

学校读者我要写书评

暂无评论

POSTER: A parallel Branch-and-Bound Algorithm with History-Based Domination 27

POSTER: A Parallel Branch-and-Bound Algorithm with History-B...

引用

27th acm sigplan symposium on principles and practice of parallel programming (PPoPP)

作者： Gonggiatgul, Taspon Shobaki, Ghassan Muyan-Ozcelik, Pinar Calif State Univ Sacramento CA USA

ISBN: (纸本)9781450392044

In this paper, we describe a parallel Branch-and-Bound (B&B) algorithm with a history-based domination technique, and we apply it to the Sequential Ordering Problem (SOP). To the best of our knowledge, the proposed algorithm is the first parallel B&B algorithm that includes a history-based domination technique and is the first parallel B&B algorithm for solving the SOP using a pure B&B approach. the proposed algorithm takes a pool-based approach and employs a collection of novel techniques that we have developed to achieve effective parallel exploration of the solution space, including parallel history domination, history table memory management, and a thread restart technique. the proposed algorithm was experimentally evaluated using the SOPLIB and TSPLIB benchmarks. the results show that using ten threads with a time limit of one hour on the medium-difficulty instances, the proposed algorithm gives a geometric-mean speedup of 19.9 on SOPLIB and 10.23 on TSPLIB, with super-linear speedups up to 65x seen on 17 instances.

关键词： parallel branch-and-bound sequential ordering problem combinatorial optimization NP-complete problems

来源：评论

学校读者我要写书评

暂无评论

POSTER: Towards OmpSs-2 and OpenACC Interoperation 27

POSTER: Towards OmpSs-2 and OpenACC Interoperation

引用

27th acm sigplan symposium on principles and practice of parallel programming (PPoPP)

作者： Korakitis, Orestis De Gonzalo, Simon Garcia Guidotti, Nicolas Barreto, Joao Pedro Monteiro, Jose C. Pena, Antonio J. Barcelona Supercomputing Ctr Barcelona Spain Univ Lisbon Inst Super Tecnico INESC ID Lisbon Portugal

ISBN: (纸本)9781450392044

the increasing demand in HPC to utilize accelerators has motivated the development of pragma-based directives to target these devices. OmpSs-2 and OpenACC are both directive-based solutions that allow application programmers to utilize accelerators. the two leverage distinct types of parallelism: task parallelism and data parallelism, respectively. Non-trivial scientific applications can benefit from both types of available parallelism. However, the combination of pragma-based models is difficult to coordinate, as both assume full control and are unaware of each other at runtime. We propose an interoperation mechanism to enable novel composability across pragma-based programming models. We study and propose a clear separation of duties and implement our approach by augmenting the OmpSs-2 programming model, compiler and runtime to support OmpSs-2 + OpenACC programming.

关键词： programming Productivity Data-flow Paradigm Runtime Scheduling Code Transformation parallelism GPU

来源：评论

学校读者我要写书评

暂无评论

Automatic Formal Verification of MPI-Based parallel Programs 11

Automatic Formal Verification of MPI-Based Parallel Programs

引用

16th acm symposium on principles and practice of parallel programming

作者： Siegel, Stephen F. Zirkel, Timothy K. Univ Delaware Verified Software Lab Dept Comp & Informat Sci Newark DE 19716 USA

ISBN: (纸本)9781450301190

the Toolkit for Accurate Scientific Software (TASS) is a suite of tools for the formal verification of MPI-based parallel programs used in computational science. TASS can verify various safety properties as well as compare two programs for functional equivalence. the TASS front end takes an integer n >= 1 and a C/MPI program, and constructs an abstract model of the program with n processes. Procedures, structs, (multi-dimensional) arrays, heap-allocated data, pointers, and pointer arithmetic are all representable in a TASS model. the model is then explored using symbolic execution and explicit state space enumeration. A number of techniques are used to reduce the time and memory consumed. A variety of realistic MPI programs have been verified with TASS, including Jacobi iteration and manager-worker type programs, and some subtle defects have been discovered. TASS is written in Java and is available from http://***/tass under the Gnu Public License.

关键词： Verification Symbolic execution MPI message-passing debugging verification

来源：评论

学校读者我要写书评

暂无评论

GPU Initiated OpenSHMEM: Correct and Eicient Intra-Kernel Networking for dGPUs 25

GPU Initiated OpenSHMEM: Correct and Eicient Intra-Kernel Ne...

引用

25th acm sigplan symposium on principles and practice of parallel programming (PPoPP)

作者： Hamidouche, Khaled LeBeane, Michael Adv Micro Devices Inc Santa Clara CA 95054 USA

ISBN: (纸本)9781450368186

Current state-of-the-art in GPU networking utilizes a host-centric, kernel-boundary communication model that reduces performance and increases code complexity. To address these concerns, recent works have explored performing network operations from within a GPU kernel itself. However, these approaches typically involve the CPU in the critical path, which leads to high latency and ineicient utilization of network and/or GPU resources. In this work, we introduce GPU Initiated OpenSHMEM (GIO), a new intra-kernel PGAS programming model and runtime that enables GPUs to communicate directly with a NIC without the intervention of the CPU. We accomplish this by exploring the GPU's coarse-grained memory model and correcting semantic mismatches when GPUs wish to directly interact with the network. GIO also reduces latency by relying on a novel template-based design to minimize the overhead of initiating a network operation. We illustrate that for structured applications like a Jacobi 2D stencil, GIO can improve application performance by up to 40% compared to traditional kernel-boundary networking. Furthermore, we demonstrate that on irregular applications like Sparse Triangular Solve (SpTS), GIO provides up to 44% improvement compared to existing intra-kernel networking schemes.

关键词： GPUs Distributed programming models RDMA networks

来源：评论

学校读者我要写书评

暂无评论

Incremental Flattening for Nested Data parallelism 19

Incremental Flattening for Nested Data Parallelism

引用

24th acm sigplan symposium on principles and practice of parallel programming (PPoPP)

作者： Henriksen, Troels thoroe, Frederik Elsman, Martin Oancea, Cosmin Univ Copenhagen Copenhagen Denmark

ISBN: (纸本)9781450362252

Compilation techniques for nested-parallel applications that can adapt to hardware and dataset characteristics are vital for unlocking the power of modern hardware. this paper proposes such a technique, which builds on flattening and is applied in the context of a functional data-parallel language. Our solution uses the degree of utilized parallelism as the driver for generating a multitude of code versions, which together cover all possible mappings of the application's regular nested parallelism to the levels of parallelism supported by the hardware. these code versions are then combined into one program by guarding them with predicates, whose threshold values are automatically tuned to hardware and dataset characteristics. Our unsupervised method-of statically clustering datasets to code versions-is different from autotuning work that typically searches for the combination of code transformations producing a single version, best suited for a specific dataset or on average for all datasets. We demonstrate-by fully integrating our technique in the repertoire of a compiler for the Futhark programming language-significant performance gains on two GPUs for three real-world applications, from the financial domain, and for six Rodinia benchmarks.

关键词： functional language parallel compilers GPGPU

来源：评论

学校读者我要写书评

暂无评论

Fine-grain parallel megabase sequence comparison with multiple heterogeneous GPUs 14

Fine-grain parallel megabase sequence comparison with multip...

引用

2014 19th acm sigplan symposium on principles and practice of parallel programming, PPoPP 2014

作者： De Sandes, Edans F.O. Miranda, Guillermo Melo, Alba C.M.A. Martorell, Xavier Ayguadé, Eduard University of Brasilia Brazil Universitat Politècnica de Catalunya Barcelona Supercomputing Center Spain

ISBN: (纸本)9781450326568

this paper proposes and evaluates a parallel strategy to execute the exact Smith-Waterman (SW) algorithm for megabase DNA sequences in heterogeneous multi-GPU platforms. In our strategy, the computation of a single huge SW matrix is spread over multiple GPUs, which communicate border elements to the neighbour, using a circular buffer mechanism that hides the communication overhead. We compared 4 pairs of human-chimpanzee homologous chromosomes using 2 different GPU environments, obtaining a performance of up to 140.36 GCUPS (Billion of cells processed per second) with 3 heterogeneous GPUS.

关键词： Graphics processing unit

来源：评论

学校读者我要写书评

暂无评论

Automatic Problem Size Sensitive Task Partitioning on Heterogeneous parallel Systems 13

Automatic Problem Size Sensitive Task Partitioning on Hetero...

引用

18th acm sigplan symposium on principles and practice of parallel programming

作者： Grasso, Ivan Kofler, Klaus Cosenza, Biagio Fahringer, thomas Univ Innsbruck Inst Informat A-6020 Innsbruck Austria

ISBN: (纸本)9781450319225

In this paper we propose a novel approach which automatizes task partitioning in heterogeneous systems. Our framework is based on the Insieme Compiler and Runtime infrastructure [1]. the compiler translates a single-device OpenCL program into a multi-device OpenCL program. the runtime system then performs dynamic task partitioning based on an offline-generated prediction model. In order to derive the prediction model, we use a machine learning approach that incorporates static program features as well as dynamic, input sensitive features. Our approach has been evaluated over a suite of 23 programs and achieves performance improvements compared to an execution of the benchmarks on a single CPU and a single GPU only.

关键词： Languages Algorithms Performance heterogeneous computing compilers GPU task partitioning code analysis machine learning runtime system

来源：评论

学校读者我要写书评

暂无评论

没有更多数据了...

全选清除本页清除全部题录导出标记到“检索档案”

共34页 << < 9 10 11 12 13 14 15 16 17 18 > >>

检索报告对象比较合并检索0

隐藏清空

合并搜索

回到顶部

执行限定条件

内容：

评分：

请选择保存的检索档案：

请选择收藏分类：

订阅名称：

通借通还

温馨提示：

图书名称：

借书校区：

取书校区：

手机号码：

邮箱地址：

一卡通帐号：

电话和邮箱必须正确填写，我们会与您联系确认。

联系人：

所在院系：

联系邮箱：

联系电话：

内蒙古自治区呼和浩特市赛罕区大学西街235号邮编: 010021

建议与咨询 留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

分类表

所选分类

限定检索结果

文献类型

馆藏范围

日期分布

学科分类号

主题

机构

作者

语言

请选择保存的检索档案： 新增检索档案 确定 取消

请选择收藏分类： 新增自定义分类 确定 取消

通借通还

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

请选择保存的检索档案：

请选择收藏分类：