image-to-image translation tasks which have been widely investigated with generative adversarial networks (GAN) aim to map an image from the source domain to the target domain. The translated image can be inversely ma...
详细信息
ISBN:
(纸本)9781728185514
image-to-image translation tasks which have been widely investigated with generative adversarial networks (GAN) aim to map an image from the source domain to the target domain. The translated image can be inversely mapped to the reconstructed source image. However, existing GAN-based schemes lack the ability to accomplish reversible translation. To remedy this drawback, a nearly reversible image-to-image translation scheme where the reconstructed source image is approximately distortion-free compared with the corresponding source image is proposed in this paper. The proposed scheme jointly considers inter-frame coding and embedding. Firstly, we organize the GAN-generated reconstructed source image and the source image into a pseudo video. Furthermore, the bitstream obtained by inter-frame coding is reversibly embedded in the translated image for nearly lossless source image reconstruction. Extensive experimental results and analysis demonstrate that the proposed scheme can achieve a high level of performance in image quality and security.
Nowadays, Typeface plays an increasingly important role in dynamic digital interfaces, but there still has little direct evaluation of visualimage perception related to the typeface design, especially for the use in ...
详细信息
ISBN:
(纸本)9781665424257
Nowadays, Typeface plays an increasingly important role in dynamic digital interfaces, but there still has little direct evaluation of visualimage perception related to the typeface design, especially for the use in interface typography. The research is based on the analysis of display screen, elaborates upon the connection between display resolution and typeface design, the relationship between display polarity and the principle of vision optics. Furthermore, essential attributes and requirements of the two genre of interface font are inspected from the human visualimage perception. Additionally, the visualprocessing of text information and visual characteristics in scanning state are elaborated, visual Angle and spatial frequency of visual perception are identified as the cornerstones influencing the design of a typeface for user interface. The methodology of visual perception can be adapted to investigate questions relevant to typographic and typeface design.
In recent years, deep learning has achieved significant progress in many respects. However, unlike other research fields with millions of labeled data such as image recognition, only several thousand labeled images ar...
详细信息
ISBN:
(纸本)9781728185514
In recent years, deep learning has achieved significant progress in many respects. However, unlike other research fields with millions of labeled data such as image recognition, only several thousand labeled images are available in image quality assessment (IQA) field for deep learning, which heavily hinders the development and application for IQA. To tackle this problem, in this paper, we proposed an error self-learning semi-supervised method for no-reference (NR) IQA (ESSIQA), which is based on deep learning. We employed an advanced full reference (FR) IQA method to expand databases and supervise the training of network. In addition, the network outputs of expanding images were used as proxy labels replacing errors between subjective scores and objective scores to achieve error self-learning. Two weights of error back propagation were designed to reduce the impact of inaccurate outputs. The experimental results show that the proposed method yielded comparative effect.
With the rapid development of multi-sensor fusion technology in various industrial fields, many composite images closely related to human life have been produced. To meet the rapidly growing needs of various image-bas...
详细信息
ISBN:
(纸本)9781665475921
With the rapid development of multi-sensor fusion technology in various industrial fields, many composite images closely related to human life have been produced. To meet the rapidly growing needs of various image-based applications, we have established the first multi-source composite image (MSCI) database for image quality assessment (IQA). Our MSCI database contains 80 reference images and 1600 distorted images, generated by four advanced compression standards with five distortion levels. In particular, these five distortion levels are determined based on the first five just noticeable difference (JND) levels. Moreover, we verify the IQA performance of some representative methods on our MSCI database. The experimental results show that the performance of the existing methods on the MSCI database needs to be further improved.
Increasing the spatial resolution and frame rate of a video simultaneously has attracted attention in recent years. The current one-stage space-time video super-resolution (STVSR) methods are difficult to deal with la...
详细信息
ISBN:
(纸本)9781728185514
Increasing the spatial resolution and frame rate of a video simultaneously has attracted attention in recent years. The current one-stage space-time video super-resolution (STVSR) methods are difficult to deal with large motion and complex scenes, and are time-consuming and memory intensive. We propose an efficient STVSR framework, which can correctly handle complicated scenes such as occlusion and large motion and generate results with clearer texture. In REDS dataset, our method outperforms all existing one-stage methods. Our method is lightweight and can generate 720p frames at 16fps on a NVIDIA GTX 1080 Ti GPU.
In this paper we address visualcommunications via print-and-scan channels from an information-theoretic point of view as communications with side information that targets quality enhancement of visual data at the out...
详细信息
In this paper we address visualcommunications via print-and-scan channels from an information-theoretic point of view as communications with side information that targets quality enhancement of visual data at the output of this type of channels. The solution to this problem addresses important aspects of multimedia data processing and management. A practical approach to side information communications for printed documents based on Wyner-Ziv and Gray setups is analyzed in the paper that assumes two separated communications channels where an appropriate distributed coding should be elaborated. The printing channel is considered to be a direct visual channel for images ("analog" channel with degradations). The "digital channel" is considered to be an appropriate auxiliary channel exploited to communicate the information exploited for quality enhancement of printed-and-scanned image. We demonstrate both theoretically and practically how one can benefit from this sort of "distributed paper communications". (c) 2006 Elsevier B.V. All rights reserved.
Pixel-wise image quality assessment (IQA) algorithms, such as mean square error (MSE), mean absolute error (MAE) and peak signal-to-noise ratio (PSNR) correlate well with perceptual quality when dealing with images sh...
详细信息
ISBN:
(纸本)9781728180687
Pixel-wise image quality assessment (IQA) algorithms, such as mean square error (MSE), mean absolute error (MAE) and peak signal-to-noise ratio (PSNR) correlate well with perceptual quality when dealing with images sharing the same distortion type but not well when processingimages in different distortion types, which is inconsistent with human visual system (HVS). Although a large number of metrics based on image error has been proposed, there are still difficulties and limitations. To solve this problem, a full reference image quality assessment (FR-IQA) method based on MAE is proposed in this paper. The metric divides the image error (difference between distorted image and reference image) map into smooth region and texture-edge region, calculates their mean values respectively, and then gives them different weights considering the masking effect. The key innovation of this paper is to propose a distortion significance measurement, which is a visual quality coefficient that can effectively indicate the influence of different distortion types on perceptual quality and unify them with HVS. The segmented image error maps are weighted by the distortion significance coefficient. The experimental results on four largest benchmark databases show that the most of the distortions are successfully evaluated and the results are consistent with HVS.
This paper presents a video watermarking scheme for copyright notification and protection. Each video frame is clustered into high motion and low motion regions. Dual watermarks are designed by exploiting video motion...
详细信息
ISBN:
(纸本)0819444111
This paper presents a video watermarking scheme for copyright notification and protection. Each video frame is clustered into high motion and low motion regions. Dual watermarks are designed by exploiting video motion, perceptual characteristics, and some other informative cues. They are embedded into certain macroblocks in these two regions respectively. The watermark in the low motion region is resilient to frame-level attacks such as frame dropping, frame averaging and frame reshuffling. The watermark in the high motion region is robust to statistical attacks. The detection does not need the original frame or any a priori index information of frames in video sequence. Good invisibility as well as strong robustness can be obtained. The experimental results validate the effectiveness of the system.
VCIP 2022 "Tire pattern image classification based on lightweight network challenge" aims to design lightweight networks that correctly classify tire surface tread patterns and indentation images using less ...
详细信息
ISBN:
(纸本)9781665475921
VCIP 2022 "Tire pattern image classification based on lightweight network challenge" aims to design lightweight networks that correctly classify tire surface tread patterns and indentation images using less overhead. To this end, we present a novel lightweight tire tread classification network. Concretely, we adopt the ShuffleNet-V2-x0.5 network as our backbone. To reduce the computation complexity, we introduce the Space-To-Depth and Anti-Alias Downsampling modules to pre-process the input image. Moreover, to enhance the classification ability of our model, we adopt the knowledge distillation strategy by considering Vision Transformer as the teacher network. To ensure the robustness of our model, we pre-train it on imageNet and fine-tune the training set of the challenge. Experiments on the challenge dataset demonstrate that our model achieves superior performance, with 99.00% classification accuracy, 25.51M FLOPs, and 0.20M parameters.
Recently, deep learning-based video compression algorithms have achieved competitive performance in Bjontegaard delta (BD) rate, especially those adopting super-resolution networks as post-processing modules in downsa...
详细信息
ISBN:
(纸本)9781665475921
Recently, deep learning-based video compression algorithms have achieved competitive performance in Bjontegaard delta (BD) rate, especially those adopting super-resolution networks as post-processing modules in downsampling-based video compression (DBC) frameworks. However, limited by the non-differentiable characteristics of traditional codecs, DBC frameworks mainly focus on improving the performance of super-resolution modules while ignoring optimizing downscaling modules. It is crucial to improve video compression performance without introducing additional modifications to the decoder client in practical application scenarios. We propose a context-aware processing network (CPN) compatible with standard codecs with no computational burden introduced to the client, which preserves the critical information and essential structures during downscaling. The proposed CPN works as a precoder cascaded by standard codecs to improve the compression performance on the server before encoding and transmission. Besides, a surrogate codec is employed to simulate the degradation process of the standard codecs and backpropagate the gradient to optimize the CPN. Experimental results show that the proposed method outperforms latest pre-processing networks and achieves considerable performance compared with the latest DBC frameworks.
暂无评论