Fast and High-Performance Learned Image Compression With Improved Checkerboard Context Model, Deformable Residual Module, and Knowledge Distillation.

Haisheng Fu Feng Liang Jie Liang Yongqiang Wang Zhenman Fang Guohe Zhang Jingning Han

IEEE Trans Image Process

Published: August 2024

Deep learning-based image compression has made great progresses recently. However, some leading schemes use serial context-adaptive entropy model to improve the rate-distortion (R-D) performance, which is very slow. In addition, the complexities of the encoding and decoding networks are quite high and not suitable for many practical applications. In this paper, we propose four techniques to balance the trade-off between the complexity and performance. We first introduce the deformable residual module to remove more redundancies in the input image, thereby enhancing compression performance. Second, we design an improved checkerboard context model with two separate distribution parameter estimation networks and different probability models, which enables parallel decoding without sacrificing the performance compared to the sequential context-adaptive model. Third, we develop a three-pass knowledge distillation scheme to retrain the decoder and entropy coding, and reduce the complexity of the core decoder network, which transfers both the final and intermediate results of the teacher network to the student network to improve its performance. Fourth, we introduce L regularization to make the numerical values of the latent representation more sparse, and we only encode non-zero channels in the encoding and decoding process to reduce the bit rate. This also reduces the encoding and decoding time. Experiments show that compared to the state-of-the-art learned image coding scheme, our method can be about 20 times faster in encoding and 70-90 times faster in decoding, and our R-D performance is also 2.3% higher. Our method achieves better rate-distortion performance than classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods, as measured by both PSNR and MS-SSIM metrics on the Kodak and Tecnick-40 datasets.

Download full-text PDF	Source
http://dx.doi.org/10.1109/TIP.2024.3445737	DOI Listing

Publication Analysis

Top Keywords

encoding decoding

learned image

image compression

improved checkerboard

checkerboard context

context model

deformable residual

residual module

knowledge distillation

r-d performance

Similar Publications

Lightweight Retinal Layer Segmentation With Global Reasoning.

IEEE Trans Instrum Meas

May 2024

School of Mechanical Engineering, Shandong University, Jinan 250061, Shandong, China.

Xiang He Weiye Song Yiming Wang Fabio Poiesi Ji Yi

Automatic retinal layer segmentation with medical images, such as optical coherence tomography (OCT) images, serves as an important tool for diagnosing ophthalmic diseases. However, it is challenging to achieve accurate segmentation due to low contrast and blood flow noises presented in the images. In addition, the algorithm should be light-weight to be deployed for practical clinical applications.

View Article and Find Full Text PDF

Similar Publications

Improving auditory attention decoding by classifying intracranial responses to glimpsed and masked acoustic events.

Imaging Neurosci (Camb)

April 2024

Department of Electrical Engineering, Columbia University, New York, NY, United States.

Vinay S Raghavan James O'Sullivan Jose Herrero Stephan Bickel Ashesh D Mehta

Listeners with hearing loss have trouble following a conversation in multitalker environments. While modern hearing aids can generally amplify speech, these devices are unable to tune into a target speaker without first knowing to which speaker a user aims to attend. Brain-controlled hearing aids have been proposed using auditory attention decoding (AAD) methods, but current methods use the same model to compare the speech stimulus and neural response, regardless of the dynamic overlap between talkers which is known to influence neural encoding.

View Article and Find Full Text PDF

Similar Publications

Emotion analysis of EEG signals using proximity-conserving auto-encoder (PCAE) and ensemble techniques.

Cogn Neurodyn

December 2025

Department of Computer Science and Engineering, Sathyabama Institute of Science and Technology, Chennai, TamilNadu India.

R Mathumitha A Maryposonia

Emotion recognition plays a crucial role in brain-computer interfaces (BCI) which helps to identify and classify human emotions as positive, negative, and neutral. Emotion analysis in BCI maintains a substantial perspective in distinct fields such as healthcare, education, gaming, and human-computer interaction. In healthcare, emotion analysis based on electroencephalography (EEG) signals is deployed to provide personalized support for patients with autism or mood disorders.

View Article and Find Full Text PDF

Similar Publications

TCN-Transformer Deep Network with Random Forest for Prediction of the Chemical Synthetic Ammonia Process.

ACS Omega

January 2025

School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China.

Jianguo Dong Xiaona Liu Ruixian Su Huimin Xu Tianyu Yu

It is of great significance to realize the accurate prediction of the key output response of the chemical synthetic ammonia process for optimizing system performance and operation monitoring. Because many key intermediate variables of complex systems are difficult to measure comprehensively, there are great difficulties and errors in mechanism analysis and identification modeling techniques. Based on random forest (RF) variable selection, a deep neural network combining temporal convolutional network (TCN) and transformer is proposed to predict the output variables of the synthetic ammonia process.

View Article and Find Full Text PDF

Similar Publications

LiteMamba-Bound: A Lightweight Mamba-based Model with Boundary-Aware and Normalized Active Contour Loss for Skin Lesion Segmentation.

Methods

January 2025

School of Electrical and Electronic Engineering, Hanoi University of Science and Technology, Hanoi, Viet Nam. Electronic address:

Quang-Huy Ho Thi-Nhu-Quynh Nguyen Thi-Thao Tran Van-Truong Pham

In the field of medical science, skin segmentation has gained significant importance, particularly in dermatology and skin cancer research. This domain demands high precision in distinguishing critical regions (such as lesions or moles) from healthy skin in medical images. With growing technological advancements, deep learning models have emerged as indispensable tools in addressing these challenges.

View Article and Find Full Text PDF

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!