检索结果-内蒙古大学图书馆

您好，读者！请登录

内蒙古大学图书馆

首页
概况
党建
资源
服务
科研支持
- 论文收录引用证明
- 科技查新
知识产权
档案馆
帮助

咨询与建议

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

您的常用邮箱：*

您的手机号码：*

问题描述：

当前已输入0个字，您还可以输入200个字

全部搜索
期刊论文
图书
学位论文
标准
纸本馆藏
外文资源发现
数据库导航
超星发现

高级检索

分类表

所选分类

>> <<

限定检索结果

标题

标题
作者
主题词
出版物名称
出版社
机构
学科分类号
摘要
ISBN
ISSN
基金资助
索书号

作者

作者
标题
主题词
出版物名称
出版社
机构
学科分类号
摘要
ISBN
ISSN
基金资助
索书号

文献类型

50,636 篇 会议
1,423 册 图书
1,044 篇 期刊文献
1 篇 学位论文

馆藏范围

53,101 篇 电子文献
3 种 纸本馆藏

日期分布

学科分类号

31,927 篇 工学
- 24,897 篇 计算机科学与技术...
- 12,629 篇 软件工程
- 5,176 篇 光学工程
- 4,760 篇 电气工程
- 4,463 篇 信息与通信工程
- 4,261 篇 机械工程
- 3,980 篇 控制科学与工程
- 2,477 篇 生物工程
- 1,736 篇 生物医学工程（可授...
- 1,583 篇 仪器科学与技术
- 1,314 篇 电子科学与技术（可...
- 795 篇 化学工程与技术
- 715 篇 安全科学与工程
- 560 篇 交通运输工程
- 383 篇 建筑学
- 335 篇 土木工程
11,899 篇 理学
- 6,481 篇 物理学
- 5,426 篇 数学
- 2,765 篇 生物学
- 1,915 篇 统计学（可授理学、...
- 804 篇 化学
- 669 篇 系统科学
5,313 篇 医学
- 5,103 篇 临床医学
- 731 篇 基础医学(可授医学...
- 459 篇 药学(可授医学、理...
3,369 篇 管理学
- 1,964 篇 图书情报与档案管...
- 1,554 篇 管理科学与工程(可...
- 485 篇 工商管理
720 篇 艺术学
- 718 篇 设计学（可授艺术学...
434 篇 法学
- 406 篇 社会学
302 篇 农学
198 篇 教育学
166 篇 经济学
63 篇 文学
48 篇 军事学

主题

17,404 篇 computer vision
9,026 篇 pattern recognit...
4,196 篇 training
3,830 篇 feature extracti...
3,134 篇 cameras
2,876 篇 computational mo...
2,794 篇 image segmentati...
2,622 篇 visualization
2,574 篇 shape
2,535 篇 face recognition
2,176 篇 robustness
2,124 篇 computer science
1,975 篇 object detection
1,960 篇 computer archite...
1,882 篇 layout
1,853 篇 object recogniti...
1,801 篇 three-dimensiona...
1,725 篇 neural networks
1,705 篇 humans
1,697 篇 image recognitio...

机构

165 篇 univ chinese aca...
144 篇 tsinghua univers...
135 篇 national laborat...
106 篇 univ sci & techn...
104 篇 zhejiang univers...
101 篇 shanghai jiao to...
95 篇 university of sc...
95 篇 microsoft resear...
85 篇 zhejiang univ pe...
84 篇 shanghai ai lab ...
74 篇 school of comput...
69 篇 computer vision ...
68 篇 peking univ peop...
68 篇 chinese acad sci...
66 篇 chinese univ hon...
63 篇 institute of inf...
62 篇 google res mount...
61 篇 univ oxford oxfo...
59 篇 univ toronto on
57 篇 swiss fed inst t...

作者

92 篇 van gool luc
87 篇 umapada pal
78 篇 zhang lei
64 篇 lee seong-whan
50 篇 vittorio murino
42 篇 yang yi
34 篇 nassir navab
34 篇 ling haibin
33 篇 li xin
33 篇 jie yang
32 篇 liu yang
31 篇 loy chen change
30 篇 escalera sergio
30 篇 h. bischof
29 篇 zhou jie
29 篇 vasconcelos nuno
29 篇 jan-michael frah...
28 篇 blumenstein mich...
27 篇 jia yunde
27 篇 luo ping

语言

50,122 篇 英文
2,746 篇 其他
252 篇 中文
22 篇 土耳其文
4 篇 西班牙文
2 篇 日文
2 篇 葡萄牙文
2 篇 俄文

检索条件"任意字段=IEEE Conference on Computer Vision and Pattern Recognition"

共 53104 条记录，以下是171-180 订阅

全选清除本页清除全部题录导出标记到"检索档案"

详细简洁

排序：

相关度排序

相关度排序
时效性降序
时效性升序

GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation

GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Khanna, Mukul Ramrakhya, Ram Chhablani, Gunjan Yenamandra, Sriram Gervet, Theophile Chang, Matthew Kiraly, Zsolt Chaplot, Devendra Singh Batra, Dhruv Mottaghi, Roozbeh Georgia Inst Technol Atlanta GA 30332 USA Carnegie Mellon Univ Pittsburgh PA 15213 USA Univ Illinois Urbana IL USA Mistral AI Paris France Univ Washington Seattle WA USA

ISBN: (纸本)9798350353006

The Embodied AI community has made significant strides in visual navigation tasks, exploring targets from 3D coordinates, objects, language descriptions, and images. However, these navigation models often handle only a single input modality as the target. With the progress achieved so far, it is time to move towards universal navigation models capable of handling various goal types, enabling more effective user interaction with robots. To facilitate this goal, we propose GOAT-Bench, a benchmark for the universal navigation task referred to as GO to AnyThing (GOAT). In this task, the agent is directed to navigate to a sequence of targets specified by the category name, language description, or image in an open-vocabulary fashion. We benchmark monolithic RL and modular methods on the GOAT task, analyzing their performance across modalities, the role of explicit and implicit scene memories, their robustness to noise in goal specifications, and the impact of memory in lifelong scenarios.

关键词： computer vision Embodied AI Visual navigation

来源：评论

学校读者我要写书评

暂无评论

VLM-PL: Advanced Pseudo Labeling approach for Class Incremental Object Detection via vision-Language Model

VLM-PL: Advanced Pseudo Labeling approach for Class Incremen...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Kim, Junsu Ku, Yunhoe Kim, Jihyeon Cha, Junuk Baek, Seungryul UNIST Ulsan South Korea MODULABS Seoul South Korea

ISBN: (纸本)9798350365474

In the field of Class Incremental Object Detection (CIOD), creating models that can continuously learn like humans is a major challenge. Pseudo-labeling methods, although initially powerful, struggle with multi-scenario incremental learning due to their tendency to forget past knowledge. To overcome this, we introduce a new approach called vision-Language Model assisted Pseudo-Labeling (VLM-PL). This technique uses vision-Language Model (VLM) to verify the correctness of pseudo ground-truths (GTs) without requiring additional model training. VLM-PL starts by deriving pseudo GTs from a pre-trained detector. Then, we generate custom queries for each pseudo GT using carefully designed prompt templates that combine image and text features. This allows the VLM to classify the correctness through its responses. Furthermore, VLM-PL integrates refined pseudo and real GTs from upcoming training, effectively combining new and old knowledge. Extensive experiments conducted on the Pascal VOC and MS COCO datasets not only highlight VLM-PL's exceptional performance in multi-scenario but also illuminate its effectiveness in dual-scenario by achieving state-of-the-art results in both.

关键词： CIOD Class Incremental Object Detection Continual Learning Incremental Learning Object Detection Pseudo Labeling vision-Language Model

来源：评论

学校读者我要写书评

暂无评论

VMRNN: Integrating vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting

VMRNN: Integrating Vision Mamba and LSTM for Efficient and A...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Tang, Yujin Dong, Peijie Tang, Zhenheng Chu, Xiaowen Liang, Junwei Hong Kong Univ Sci & Technol Guangzhou AI Thrust Guangzhou Peoples R China Hong Kong Univ Sci & Technol Guangzhou DSA Thrust Guangzhou Peoples R China Hong Kong Baptist Univ Dept Comp Sci Hong Kong Peoples R China Hong Kong Univ Sci & Technol Dept Comp Sci & Engn Hong Kong Peoples R China

ISBN: (纸本)9798350365474

Combining Convolutional Neural Networks (CNNs) or vision Transformers(ViTs) with Recurrent Neural Networks (RNNs) for spatiotemporal forecasting has yielded unparalleled results in predicting temporal and spatial dynamics. However, modeling extensive global information remains a formidable challenge;CNNs are limited by their narrow receptive fields, and ViTs struggle with the intensive computational demands of their attention mechanisms. The emergence of recent Mamba-based architectures has been met with enthusiasm for their exceptional long-sequence modeling capabilities, surpassing established vision models in efficiency and accuracy, which motivates us to develop an innovative architecture tailored for spatiotemporal forecasting. In this paper, we propose the VMRNN cell, a new recurrent unit that integrates the strengths of vision Mamba blocks with LSTM. We construct a network centered on VMRNN cells to tackle spatiotemporal prediction tasks effectively. Our extensive evaluations show that our proposed approach secures competitive results on a variety of tasks while maintaining a smaller model size. Our code is available at https://***/yyyujintang/VMRNN-PyTorch.

关键词： Spatiotemporal Forecasting State Space Model Video Prediction

来源：评论

学校读者我要写书评

暂无评论

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

Summarize the Past to Predict the Future: Natural Language D...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Pasca, Razvan-George Gavryushin, Alexey Hamza, Muhammad Kuo, Yen-Ling Mo, Kaichun Van Gool, Luc Hilliges, Otmar Wang, Xi Swiss Fed Inst Technol Zurich Switzerland Univ Zurich Zurich Switzerland Univ Virginia Charlottesville VA USA NVIDIA Santa Clara CA USA Katholieke Univ Leuven Leuven Belgium INSAIT Sofia Bulgaria

ISBN: (纸本)9798350353006

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects, coined action context. We propose TransFusion, a multimodal transformer-based architecture for short-term object interaction anticipation. Our method exploits the representational power of language by summarizing the action context textually, after leveraging pre-trained vision-language foundation models to extract the action context from past video frames. The summarized action context and the last observed video frame are processed by the multimodal fusion module to forecast the next object interaction. Experiments on the Ego4D next active object interaction dataset show the effectiveness of our multimodal fusion model and highlight the benefits of using the power of foundation models and language-based context summaries in a task where vision may appear to suffice. Our novel approach outperforms all state-of-the-art methods on both versions of the Ego4D dataset. A project video and code are available at https://***/transfusion-proj/.

关键词： computer vision Egocentric vision Human Behavior Prediction Multimodal Learning Object Interaction Anticipation vision-Language Models

来源：评论

学校读者我要写书评

暂无评论

CUE-Net: Violence Detection Video Analytics with Spatial Cropping, Enhanced UniformerV2 and Modified Efficient Additive Attention

CUE-Net: Violence Detection Video Analytics with Spatial Cro...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Senadeera, Damith Chamalke Yang, Xiaoyun Kollias, Dimitrios Slabaugh, Gregory Queen Mary Univ London Sch Elect Engn & Comp Sci London England Queen Marys Digital Environm Res Inst DERI London England Remark AI UK Ltd London England

ISBN: (纸本)9798350365474

In this paper we introduce CUE-Net, a novel architecture designed for automated violence detection in video surveillance. As surveillance systems become more prevalent due to technological advances and decreasing costs, the challenge of efficiently monitoring vast amounts of video data has intensified. CUE-Net addresses this challenge by combining spatial Cropping with an enhanced version of the UniformerV2 architecture, integrating convolutional and self-attention mechanisms alongside a novel Modified Efficient Additive Attention mechanism (which reduces the quadratic time complexity of self-attention) to effectively and efficiently identify violent activities. This approach aims to overcome traditional challenges such as capturing distant or partially obscured subjects within video frames. By focusing on both local and global spatio-temporal features, CUE-Net achieves state-of-the-art performance on the RWF-2000 and RLVS datasets, surpassing existing methods. The source code is available at (1).

关键词： computer vision Cropping Deep Learning Efficient Additive Attention UniFormerV2 Video Analytics Violence Detection

来源：评论

学校读者我要写书评

暂无评论

How Much You Ate? Food Portion Estimation on Spoons

How Much You Ate? Food Portion Estimation on Spoons

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Sharma, Aaryam Czarnecki, Chris Chen, Yuhao Xi, Pengcheng Xu, Linlin Wong, Alexander Univ Waterloo Vis & Image Proc Lab Waterloo ON Canada Natl Res Council Canada Ottawa ON Canada

ISBN: (纸本)9798350365474

Monitoring dietary intake is a crucial aspect of promoting healthy living. In recent years, advances in computer vision technology have facilitated dietary intake monitoring through the use of images and depth cameras. However, the current state-of-the-art image-based food portion estimation algorithms assume that users take images of their meals one or two times, which can be inconvenient and fail to capture food items that are not visible from a top-down perspective, such as ingredients submerged in a stew. To address these limitations, we introduce an innovative solution that utilizes stationary user-facing cameras to track food items on utensils, not requiring any change of camera perspective after installation. The shallow depth of utensils provides a more favorable angle for capturing food items, and tracking them on the utensil's surface offers a significantly more accurate estimation of dietary intake without the need for post-meal image capture. The system is reliable for estimation of nutritional content of liquid-solid heterogeneous mixtures such as soups and stews. Through a series of experiments, we demonstrate the exceptional potential of our method as a non-invasive, user-friendly, and highly accurate dietary intake monitoring tool.

关键词： computer-vision estimation food nutrition volumetric

来源：评论

学校读者我要写书评

暂无评论

GreedyViG: Dynamic Axial Graph Construction for Efficient vision GNNs

GreedyViG: Dynamic Axial Graph Construction for Efficient Vi...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Munir, Mustafa Avery, William Rahman, Md Mostafijur Marculescu, Radu Univ Texas Austin Austin TX 78712 USA

ISBN: (纸本)9798350353013;9798350353006

vision graph neural networks (ViG) offer a new avenue for exploration in computer vision. A major bottleneck in ViGs is the inefficient k-nearest neighbor (KNN) operation used for graph construction. To solve this issue, we propose a new method for designing ViGs, Dynamic Axial Graph Construction (DAGC), which is more efficient than KNN as it limits the number of considered graph connections made within an image. Additionally, we propose a novel CNN-GNN architecture, GreedyViG, which uses DAGC. Extensive experiments show that GreedyViG beats existing ViG, CNN, and ViT architectures in terms of accuracy, GMACs, and parameters on image classification, object detection, instance segmentation, and semantic segmentation tasks. Our smallest model, GreedyViG-S, achieves 81.1% top-1 accuracy on ImageNet-1K, 2.9% higher than vision GNN and 2.2% higher than vision HyperGraph Neural Network (ViHGNN), with less GMACs and a similar number of parameters. Our largest model, GreedyViG-B obtains 83.9% top-1 accuracy, 0.2% higher than vision GNN, with a 66.6% decrease in parameters and a 69% decrease in GMACs. GreedyViG-B also obtains the same accuracy as ViHGNN with a 67.3% decrease in parameters and a 71.3% decrease in GMACs. Our work shows that hybrid CNN-GNN architectures not only provide a new avenue for de-signing efficient models, but that they can also exceed the performance of current state-of-the-art models(1).

关键词： Deep Learning Efficient computer vision Graph Neural Networks

来源：评论

学校读者我要写书评

暂无评论

Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box vision-Language Models for Selective Visual Question Answering

Consistency and Uncertainty: Identifying Unreliable Response...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Khan, Zaid Fu, Yun Northeastern Univ Boston MA 02115 USA

ISBN: (纸本)9798350353006

The goal of selective prediction is to allow an a model to abstain when it may not be able to deliver a reliable prediction, which is important in safety-critical contexts. Existing approaches to selective prediction typically require access to the internals of a model, require retraining a model or study only unimodal models. However, the most powerful models (e.g. GPT-4) are typically only available as black boxes with inaccessible internals, are not retrainable by end-users, and are frequently used for multimodal tasks. We study the possibility of selective prediction for vision-language models in a realistic, black-box setting. We propose using the principle of neighborhood consistency to identify unreliable responses from a black-box vision-language model in question answering tasks. We hypothesize that given only a visual question and model response, the consistency of the model's responses over the neighborhood of a visual question will indicate reliability. It is impossible to directly sample neighbors in feature space in a black-box setting. Instead, we show that it is possible to use a smaller proxy model to approximately sample from the neighborhood. We find that neighborhood consistency can be used to identify model responses to visual questions that are likely unreliable, even in adversarial settings or settings that are out-of-distribution to the proxy model.

关键词： predictive uncertainty selective prediction trustworthy ml vision-language visual question answering

来源：评论

学校读者我要写书评

暂无评论

QAttn: Efficient GPU Kernels for mixed-precision vision Transformers

QAttn: Efficient GPU Kernels for mixed-precision Vision Tran...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Kluska, Piotr Castello, Adrian Scheidegger, Florian Malossi, A. Cristiano I. Quintana-Orti, Enrique S. IBM Res Europe Ruschlikon Switzerland Univ Politecn Valencia Valencia Spain

ISBN: (纸本)9798350365474

vision Transformers have demonstrated outstanding performance in computer vision tasks. Nevertheless, this superior performance for large models comes at the expense of increasing memory usage for storing the parameters and intermediate activations. To accelerate model inference, in this work we develop and evaluate integer and mixed-precision kernels in Triton for the efficient execution of two fundamental building blocks of transformers -linear layer and attention- on graphics processing units (GPUs). On an NVIDIA A100 GPU, our kernel implementations of vision Transformers achieve a throughput speedup of up to 7x compared with reference kernels in PyTorch floating-point single precision (FP32). Additionally, the accuracy for the ViT Large model top-1 drops by less than one percent on the ImageNet1K classification task. We also observe up to 6x increased throughput by applying our kernels to the Segment Anything Model image encoder while keeping the mIOU close to the FP32 reference on the COCO2017 dataset for static and dynamic quantization. Furthermore, our kernels demonstrate improved speed to the TensorRT INT8 linear layer, and we improve the throughput of base FP16 (half precision) Triton attention on average by up to 19 +/- 4.01%. We have open-sourced the QAtnn framework, which is tightly integrated with the PyTorch quantization workflow https://***/IBM/qattn.

关键词： compression instance segmentation object classification quantization vision transformers

来源：评论

学校读者我要写书评

暂无评论

Building vision-Language Models on Solid Foundations with Masked Distillation

Building Vision-Language Models on Solid Foundations with Ma...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Sameni, Sepehr Kafle, Kushal Tan, Hao Jenni, Simon Univ Bern Bern Switzerland Adobe Res San Jose CA USA

ISBN: (纸本)9798350353006

Recent advancements in vision-Language Models (VLMs) have marked a significant leap in bridging the gap between computer vision and natural language processing. However, traditional VLMs, trained through contrastive learning on limited and noisy image-text pairs, often lack the spatial and linguistic understanding to generalize well to dense vision tasks or less common languages. Our approach, Solid Foundation CLIP (SF-CLIP), circumvents this issue by implicitly building on the solid visual and language understanding of foundational models trained on vast amounts of unimodal data. SF-CLIP integrates contrastive image-text pretraining with a masked knowledge distillation from large foundational text and vision models. This methodology guides our VLM in developing robust text and image representations. As a result, SF-CLIP shows exceptional zero-shot classification accuracy and enhanced image and text retrieval capabilities, setting a new state of the art for ViT-B/16 trained on YFCC15M and CC12M. Moreover, the dense per-patch supervision enhances our zero-shot and linear probe performance in semantic segmentation tasks. A remarkable aspect of our model is its multilingual proficiency, evidenced by strong retrieval results in multiple languages despite being trained predominantly on English data. We achieve all of these improvements without sacrificing the training efficiency through our selective application of masked distillation and the inheritance of teacher word embeddings.

关键词： CLIP Distillation LLM Multilingual Multimodal Representation Learning

来源：评论

学校读者我要写书评

暂无评论

没有更多数据了...

全选清除本页清除全部题录导出标记到“检索档案”

共500页 << < 14 15 16 17 18 19 20 21 22 23 > >>

检索报告对象比较合并检索0

隐藏清空

合并搜索

回到顶部

执行限定条件

内容：

评分：

请选择保存的检索档案：

请选择收藏分类：

订阅名称：

通借通还

温馨提示：

图书名称：

借书校区：

取书校区：

手机号码：

邮箱地址：

一卡通帐号：

电话和邮箱必须正确填写，我们会与您联系确认。

联系人：

所在院系：

联系邮箱：

联系电话：

内蒙古自治区呼和浩特市赛罕区大学西街235号邮编: 010021

建议与咨询 留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

分类表

所选分类

限定检索结果

文献类型

馆藏范围

日期分布

学科分类号

主题

机构

作者

语言

请选择保存的检索档案： 新增检索档案 确定 取消

请选择收藏分类： 新增自定义分类 确定 取消

通借通还

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

请选择保存的检索档案：

请选择收藏分类：