检索结果-内蒙古大学图书馆

您好，读者！请登录

内蒙古大学图书馆

首页
概况
党建
资源
服务
科研支持
- 论文收录引用证明
- 科技查新
知识产权
档案馆
帮助

咨询与建议

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

您的常用邮箱：*

您的手机号码：*

问题描述：

当前已输入0个字，您还可以输入200个字

全部搜索
期刊论文
图书
学位论文
标准
纸本馆藏
外文资源发现
数据库导航
超星发现

高级检索

分类表

所选分类

>> <<

限定检索结果

标题

标题
作者
主题词
出版物名称
出版社
机构
学科分类号
摘要
ISBN
ISSN
基金资助
索书号

作者

作者
标题
主题词
出版物名称
出版社
机构
学科分类号
摘要
ISBN
ISSN
基金资助
索书号

文献类型

50,693 篇 会议
1,424 册 图书
1,049 篇 期刊文献
1 篇 学位论文

馆藏范围

53,164 篇 电子文献
3 种 纸本馆藏

日期分布

学科分类号

31,979 篇 工学
- 24,946 篇 计算机科学与技术...
- 12,662 篇 软件工程
- 5,176 篇 光学工程
- 4,776 篇 电气工程
- 4,488 篇 信息与通信工程
- 4,263 篇 机械工程
- 4,009 篇 控制科学与工程
- 2,480 篇 生物工程
- 1,737 篇 生物医学工程（可授...
- 1,583 篇 仪器科学与技术
- 1,332 篇 电子科学与技术（可...
- 795 篇 化学工程与技术
- 729 篇 安全科学与工程
- 570 篇 交通运输工程
- 389 篇 建筑学
- 339 篇 土木工程
11,923 篇 理学
- 6,487 篇 物理学
- 5,441 篇 数学
- 2,768 篇 生物学
- 1,918 篇 统计学（可授理学、...
- 804 篇 化学
- 669 篇 系统科学
5,318 篇 医学
- 5,105 篇 临床医学
- 732 篇 基础医学(可授医学...
- 459 篇 药学(可授医学、理...
3,389 篇 管理学
- 1,977 篇 图书情报与档案管...
- 1,565 篇 管理科学与工程(可...
- 487 篇 工商管理
720 篇 艺术学
- 718 篇 设计学（可授艺术学...
439 篇 法学
- 411 篇 社会学
303 篇 农学
199 篇 教育学
167 篇 经济学
63 篇 文学
48 篇 军事学

主题

17,427 篇 computer vision
9,029 篇 pattern recognit...
4,199 篇 training
3,832 篇 feature extracti...
3,134 篇 cameras
2,879 篇 computational mo...
2,796 篇 image segmentati...
2,624 篇 visualization
2,574 篇 shape
2,536 篇 face recognition
2,176 篇 robustness
2,125 篇 computer science
1,976 篇 object detection
1,961 篇 computer archite...
1,882 篇 layout
1,855 篇 object recogniti...
1,801 篇 three-dimensiona...
1,725 篇 neural networks
1,705 篇 humans
1,700 篇 image recognitio...

机构

165 篇 univ chinese aca...
144 篇 tsinghua univers...
135 篇 national laborat...
106 篇 univ sci & techn...
104 篇 zhejiang univers...
101 篇 shanghai jiao to...
95 篇 university of sc...
95 篇 microsoft resear...
85 篇 zhejiang univ pe...
84 篇 shanghai ai lab ...
74 篇 school of comput...
69 篇 computer vision ...
68 篇 peking univ peop...
68 篇 chinese acad sci...
66 篇 chinese univ hon...
63 篇 institute of inf...
62 篇 google res mount...
61 篇 univ oxford oxfo...
59 篇 univ toronto on
57 篇 swiss fed inst t...

作者

92 篇 van gool luc
87 篇 umapada pal
78 篇 zhang lei
64 篇 lee seong-whan
50 篇 vittorio murino
42 篇 yang yi
34 篇 nassir navab
34 篇 ling haibin
33 篇 li xin
33 篇 jie yang
31 篇 loy chen change
31 篇 liu yang
30 篇 escalera sergio
30 篇 h. bischof
29 篇 zhou jie
29 篇 vasconcelos nuno
29 篇 jan-michael frah...
28 篇 blumenstein mich...
27 篇 jia yunde
27 篇 luo ping

语言

51,180 篇 英文
1,749 篇 其他
253 篇 中文
22 篇 土耳其文
4 篇 西班牙文
2 篇 日文
2 篇 葡萄牙文
2 篇 俄文
1 篇 法文

检索条件"任意字段=IEEE Conference on Computer Vision and Pattern Recognition"

共 53167 条记录，以下是281-290 订阅

全选清除本页清除全部题录导出标记到"检索档案"

详细简洁

排序：

相关度排序

相关度排序
时效性降序
时效性升序

RMT: Retentive Networks Meet vision Transformers

RMT: Retentive Networks Meet Vision Transformers

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Fan, Qihang Huang, Huaibo Chen, Mingrui Liu, Hongmin He, Ran Chinese Acad Sci Inst Automat MAIS & CRIPAC Beijing Peoples R China Univ Chinese Acad Sci Sch Artificial Intelligence Beijing Peoples R China Univ Sci & Technol Beijing Beijing Peoples R China

ISBN: (纸本)9798350353013;9798350353006

vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and bears a quadratic computational complexity, thereby constraining the applicability of ViT. To alleviate these issues, we draw inspiration from the recent Retentive Network (RetNet) in the field of NLP, and propose RMT, a strong vision backbone with explicit spatial prior for general purposes. Specifically, we extend the RetNet's temporal decay mechanism to the spatial domain, and propose a spatial decay matrix based on the Manhattan distance to introduce the explicit spatial prior to Self-Attention. Additionally, an attention decomposition form that adeptly adapts to explicit spatial prior is proposed, aiming to reduce the computational burden of modeling global information without disrupting the spatial decay matrix. Based on the spatial decay matrix and the attention decomposition form, we can flexibly integrate explicit spatial prior into the vision backbone with linear complexity. Extensive experiments demonstrate that RMT exhibits exceptional performance across various vision tasks. Specifically, without extra training data, RMT achieves 84.8% and 86.1% top-1 acc on ImageNet-1k with 27M/4.5GFLOPs and 96M/18.2GFLOPs. For downstream tasks, RMT achieves 54.5 box AP and 47.2 mask AP on the COCO detection task, and 52.8 mIoU on the ADE20K se-mantic segmentation task.

关键词： vision Transformer

来源：评论

学校读者我要写书评

暂无评论

LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation

LQMFormer: Language-aware Query Mask Transformer for Referri...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Shah, Nisarg A. Vibashan, V. S. Patel, Vishal M. Johns Hopkins Univ Baltimore MD 21218 USA

ISBN: (纸本)9798350353006

Referring Image Segmentation (RIS) aims to segment objects from an image based on a language description. Recent advancements have introduced transformer-based methods that leverage cross-modal dependencies, significantly enhancing performance in referring segmentation tasks. These methods are designed such that each query predicts different masks. However, RIS inherently requires a single-mask prediction, leading to a phenomenon known as Query Collapse, where all queries yield the same mask prediction. This reduces the generalization capability of the RIS model for complex or novel scenarios. To address this issue, we propose a Multi-modal Query Feature Fusion technique, characterized by two innovative designs: (1) Gaussian enhanced Multi-Modal Fusion, a novel visual grounding mechanism that enhances overall representation by extracting rich local visual information and global visual-linguistic relationships, and (2) A Dynamic Query Module that produces a diverse set of queries through a scoring network where the network selectively focuses on queries for objects referred to in the language description. Moreover, we show that including an auxiliary loss to increase the distance between mask representations of different queries further enhances performance and mitigates query collapse. Extensive experiments conducted on four benchmark datasets validate the effectiveness of our framework.

关键词： Image segmentation Multimodal Transformer vision-Language

来源：评论

学校读者我要写书评

暂无评论

PELA: Learning Parameter-Efficient Models with Low-Rank Approximation

PELA: Learning Parameter-Efficient Models with Low-Rank Appr...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Guo, Yangyang Wang, Guangzhi Kankanhalli, Mohan Natl Univ Singapore Singapore Singapore

ISBN: (纸本)9798350353006

Applying a pre-trained large model to downstream tasks is prohibitive under resource-constrained conditions. Re-cent dominant approaches for addressing efficiency issues involve adding a few learnable parameters to the fixed backbone model. This strategy, however, leads to more challenges in loading large models for downstream fine-tuning with limited resources. In this paper, we propose a novel method for increasing the parameter efficiency of pre-trained models by introducing an intermediate pre-training stage. To this end, we first employ low-rank approximation to compress the original large model and then devise a feature distillation module and a weight perturbation regularization module. These modules are specifically designed to enhance the low-rank model. In particular, we update only the low-rank model while freezing the backbone parameters during pre-training. This allows for direct and efficient utilization of the low-rank model for downstream fine-tuning tasks. The proposed method achieves both efficiencies in terms of required parameters and computation time while maintaining comparable results with minimal modifications to the backbone architecture. Specifically, when applied to three vision-only and one vision-language Transformer models, our approach often demonstrates a merely similar to 0.6 point decrease in performance while reducing the original parameter size by 1/3 to 2/3. We release our code at link.

关键词： Knowledge Distillation Low-rank Approximation vision-Language

来源：评论

学校读者我要写书评

暂无评论

Synthesize, Diagnose, and Optimize: Towards Fine-Grained vision-Language Understanding

Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vis...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Peng, Wujian Xi, Sicheng You, Zuyao Lan, Shiyi Wu, Zuxuan Fudan Univ Sch CS Shanghai Key Lab Intell Info Proc Shanghai Peoples R China Shanghai Collaborat Innovat Ctr Intelligent Visua Shanghai Peoples R China NVIDIA Shenzhen Guangdong Peoples R China

ISBN: (纸本)9798350353006

vision language models (VLM) have demonstrated remarkable performance across various downstream tasks. However, understanding fine-grained visual-linguistic concepts, such as attributes and inter-object relationships, remains a significant challenge. While several benchmarks aim to evaluate VLMs in finer granularity, their primary focus remains on the linguistic aspect, neglecting the visual dimension. Here, we highlight the importance of evaluating VLMs from both a textual and visual perspective. We introduce a progressive pipeline to synthesize images that vary in a specific attribute while ensuring consistency in all other aspects. Utilizing this data engine, we carefully design a benchmark, SPEC, to diagnose the comprehension of object size, position, existence, and count. Subsequently, we conduct a thorough evaluation of four leading VLMs on SPEC. Surprisingly, their performance is close to random guess, revealing significant limitations. With this in mind, we propose a simple yet effective approach to optimize VLMs in fine-grained understanding, achieving significant improvements on SPEC without compromising the zero-shot performance. Results on two additional fine-grained benchmarks also show consistent improvements, further validating the transferability of our approach. Code and data are available at https://***/wjpoom/SPEC.

关键词： Fine-grained understdanding vision language model

来源：评论

学校读者我要写书评

暂无评论

ViT-CoMer: vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions

ViT-CoMer: Vision Transformer with Convolutional Multi-scale...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Xia, Chunlong Wang, Xinliang Lv, Feng Hao, Xin Shi, Yifeng Baidu Inc Beijing Peoples R China

ISBN: (纸本)9798350353013;9798350353006

Although vision Transformer (ViT) has achieved significant success in computer vision, it does not perform well in dense prediction tasks due to the lack of inner-patch information interaction and the limited diversity of feature scale. Most existing studies are devoted to designing vision-specific transformers to solve the above problems, which introduce additional pre-training costs. Therefore, we present a plain, pre-training-free, and feature-enhanced ViT back-bone with Convolutional Multi-scale feature interaction, named ViT-CoMer, which facilitates bidirectional interaction between CNN and transformer. Compared to the state-of-the-art, ViT-CoMer has the following advantages: (1) We inject spatial pyramid multi-receptive field convolutional features into the ViT architecture, which effectively alleviates the problems of limited local information interaction and single-feature representation in ViT. (2) We propose a simple and efficient CNN-Transformer bidirectional fusion interaction module that performs multi-scale fusion across hierarchical features, which is beneficial for handling dense prediction tasks. (3) We evaluate the performance of ViT-CoMer across various dense prediction tasks, different frameworks, and multiple advanced pre-training. Notably, our ViT-CoMer-L achieves 64.3% AP on COCO val2017 without extra training data, and 62.1% mIoU on ADE20K val, both of which are comparable to state-of-the-art methods. We hope ViT-CoMer can serve as a new backbone for dense prediction tasks to facilitate future research. The code will be released at https://***/Traffic-X/ViT-CoMer.

关键词： DensePrediction visionFoundationBackbone visionTransformer

来源：评论

学校读者我要写书评

暂无评论

Video Based Computational Coding of Movement Anomalies in ASD Children

Video Based Computational Coding of Movement Anomalies in AS...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Singh, Priya Pathak, Abhishek Ganai, Umer Jon Bhushan, Braj Subramanian, Venkatesh K. TCS Res Mumbai Maharashtra India Indian Inst Technol Kanpur Kanpur India Airbase Labs India Pvt Ltd Bengaluru India

ISBN: (纸本)9798350365474

Autism spectrum disorder (ASD) is a neurodevelopmental disorder. Early detection and diagnosis are instrumental in early intervention, yet diagnosis often remains delayed due to the limited availability of clinical practitioners and specialists. We propose a computer vision and Machine Learning based novel framework for quantitative screening of autism spectrum disorder (ASD). This is aimed to minimize the need for trained professionals at the initial stage but not substitute for it. We designed simple activities in consultation with ASD clinical psychologists and therapists for children in the 3-7 years age group that could be performed in their natural environment (home). The temporal features extracted from these activities encode the behavioral differences between Autism Spectrum Disorder (ASD) and Typically Developing (TD) control groups. Due to the unavailability of a public dataset of children performing the designed task, we created our own video dataset of 210 videos taken in unconstrained natural settings. The dataset was collected from a single RGB camera. The proposed vision and learning-based algorithms extract features from the collected data for a comprehensive set of indicators including the visual attention span, name-calling response, neck pose of the subjects, gross motor movement and establish a parametrized automated protocol for early detection without the need to take the subjects out of their natural daily environment. This forestalls the possibility of misperformance by the subject out of nervousness due to unfamiliar surroundings. Results show that our ASD screening methodology can achieve superior performance compared to the single phenotype approaches, and thus has a prognostic value that could be helpful for both clinical and research applications.

关键词： Autism Spectrum Disorder computer vision Early detection of ASD Machine Learning

来源：评论

学校读者我要写书评

暂无评论

On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?

On the test-time zero-shot generalization of vision-language...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Zanella, Maxime Ben Ayed, Ismail UCLouvain Louvain Belgium UMons Mons Belgium ETS Montreal Montreal PQ Canada

ISBN: (纸本)9798350353006

The development of large vision-language models, notably CLIP, has catalyzed research into effective adaptation techniques, with a particular focus on soft prompt tuning. Conjointly, test-time augmentation, which utilizes multiple augmented views of a single image to enhance zero-shot generalization, is emerging as a significant area of interest. This has predominantly directed research efforts toward test-time prompt tuning. In contrast, we introduce a robust MeanShift for Test-time Augmentation (MTA), which surpasses prompt-based methods without requiring this intensive training procedure. This positions MTA as an ideal solution for both standalone and API-based applications. Additionally, our method does not rely on ad hoc rules (e.g., confidence threshold) used in some previous test-time augmentation techniques to filter the augmented views. Instead, MTA incorporates a quality assessment variable for each view directly into its optimization process, termed as the inlierness score. This score is jointly optimized with a density mode seeking process, leading to an efficient training- and hyperparameter-free approach. We extensively benchmark our method on 15 datasets and demonstrate MTA's superiority and computational efficiency. Deployed easily as plug-and-play module on top of zero-shot models and state-of-the-art few-shot methods, MTA shows systematic and consistent improvements.

关键词： CLIP test-time augmentation training-free vision-language zero-shot

来源：评论

学校读者我要写书评

暂无评论

Compositional Chain-of-Thought Prompting for Large Multimodal Models

Compositional Chain-of-Thought Prompting for Large Multimoda...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Mitra, Chancharik Huang, Brandon Darrell, Trevor Herzig, Roei Univ Calif Berkeley Berkeley CA 94720 USA

ISBN: (纸本)9798350353006

The combination of strong visual backbones and Large Language Model (LLM) reasoning has led to Large Multimodal Models (LMMs) becoming the current standard for a wide range of vision and language (VL) tasks. However, recent research has shown that even the most advanced LMMs still struggle to capture aspects of compositional visual reasoning, such as attributes and relationships between objects. One solution is to utilize scene graphs (SGs)-a formalization of objects and their relations and attributes that has been extensively used as a bridge between the visual and textual domains. Yet, scene graph data requires scene graph annotations, which are expensive to collect and thus not easily scalable. Moreover, finetuning an LMM based on SG data can lead to catastrophic forgetting of the pretraining objective. To overcome this, inspired by chain-of-thought methods, we propose Compositional Chain-of-Thought (CCoT), a novel zero-shot Chain-of-Thought prompting method that utilizes SG representations in order to extract compositional knowledge from an LMM. Specifically, we first generate an SG using the LMM, and then use that SG in the prompt to produce a response. Through extensive experiments, we find that the proposed CCoT approach not only improves LMM performance on several vision and language (VL) compositional benchmarks but also improves the performance of several popular LMMs on general multimodal benchmarks, without the need for fine-tuning or annotated ground-truth SGs. Code: https://***/chancharikmitra/CCoT.

关键词： Compositionality Large Multimodal Models Multimodality Prompting Scene Graphs vision & Language

来源：评论

学校读者我要写书评

暂无评论

ShapeWalk: Compositional Shape Editing through Language-Guided Chains

ShapeWalk: Compositional Shape Editing through Language-Guid...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Slim, Habib Elhoseiny, Mohamed KAUST Thuwal Saudi Arabia

ISBN: (纸本)9798350353006

Editing 3D shapes through natural language instructions is a challenging task that requires the comprehension of both language semantics and fine-grained geometric details. To bridge this gap, we introduce ShapeWalk, a carefully designed synthetic dataset designed to advance the field of language-guided shape editing. The dataset consists of 158K unique shapes connected through 26K edit chains, with an average length of 14 chained shapes. Each consecutive pair of shapes is associated with precise language instructions describing the applied edits. We synthesize edit chains by reconstructing and interpolating shapes sampled from a realistic CAD-designed 3D dataset in the parameter space of the GeoCode shape program. We leverage rule-based methods and language models to generate accurate and realistic natural language prompts corresponding to each edit. To illustrate the practicality of our contribution, we train neural editor modules in the latent space of shape autoencoders, and demonstrate the ability of our dataset to enable a variety of language-guided shape edits. Finally, we introduce multi-step editing metrics to benchmark the capacity of our models to perform recursive shape edits. We hope that our work will enable further study of compositional language-guided shape editing, and finds application in 3D CAD design and interactive modeling.

关键词： 3D language editing 3D vision compositionality

来源：评论

学校读者我要写书评

暂无评论

ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangling and Self-Prompting

ZePT: Zero-Shot Pan-Tumor Segmentation via Query-Disentangli...

引用

ieee/CVF conference on computer vision and pattern recognition (CVPR)

作者： Jiang, Yankai Huang, Zhongzhen Zhang, Rongzhao Zhang, Xiaofan Zhang, Shaoting Shanghai AI Lab Shanghai Peoples R China Shanghai Jiao Tong Univ Shanghai Peoples R China SenseTime Res Shanghai Peoples R China

ISBN: (纸本)9798350353006

The long-tailed distribution problem in medical image analysis reflects a high prevalence of common conditions and a low prevalence of rare ones, which poses a significant challenge in developing a unified model capable of identifying rare or novel tumor categories not encountered during training. In this paper, we propose a new Zero-shot Pan-Tumor segmentation framework (ZePT) based on query-disentangling and self-prompting to segment unseen tumor categories beyond the training set. ZePT disentangles the object queries into two subsets and trains them in two stages. Initially, it learns a set of fundamental queries for organ segmentation through an object-aware feature grouping strategy, which gathers organ-level visual features. Subsequently, it refines the other set of advanced queries that focus on the auto-generated visual prompts for unseen tumor segmentation. Moreover, we introduce query-knowledge alignment at the feature level to enhance each query's discriminative representation and generalizability. Extensive experiments on various tumor segmentation tasks demonstrate the performance superiority of ZePT, which surpasses the previous counterparts and evidences the promising ability for zero-shot tumor segmentation in real-world settings.

关键词： Medical Image Segmentation vision-Language Model

来源：评论

学校读者我要写书评

暂无评论

没有更多数据了...

全选清除本页清除全部题录导出标记到“检索档案”

共500页 << < 25 26 27 28 29 30 31 32 33 34 > >>

检索报告对象比较合并检索0

隐藏清空

合并搜索

回到顶部

执行限定条件

内容：

评分：

请选择保存的检索档案：

请选择收藏分类：

订阅名称：

通借通还

温馨提示：

图书名称：

借书校区：

取书校区：

手机号码：

邮箱地址：

一卡通帐号：

电话和邮箱必须正确填写，我们会与您联系确认。

联系人：

所在院系：

联系邮箱：

联系电话：

内蒙古自治区呼和浩特市赛罕区大学西街235号邮编: 010021

建议与咨询 留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

分类表

所选分类

限定检索结果

文献类型

馆藏范围

日期分布

学科分类号

主题

机构

作者

语言

请选择保存的检索档案： 新增检索档案 确定 取消

请选择收藏分类： 新增自定义分类 确定 取消

通借通还

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

请选择保存的检索档案：

请选择收藏分类：