NeurIPS 2024PastEfficiency
Workshop on Machine Learning and Compression, NeurIPS 2024
Compression Workshop @ NeurIPS 2024
- Submission deadline
- Oct 1, 2024, 11:59 UTCimported from OpenReview — check the website for extensions
- Submission portal
- OpenReview
- Notes
- Topics were auto-suggested and may be imprecise — edits welcome.
Accepted papers (95)
Fetched from OpenReview (v2) on 2026-06-10.
A Theory for Compressibility of Graph Transformers for Transductive Learning
A Tighter Complexity Analysis of SparseGPT
Accelerating Memory-Efficient LLM Training and Fine-Tuning via Tracking the Gradient Subspace
Adapting Language Models via Token Translation
Adaptive Quantization and Pruning of Deep Neural Networks via Layer Importance Estimation
AdaQuantLM: LLM Quantization with Adaptive Bit-Widths
An image to tailor: I-Frame Domain Adaptation in Neural Video Compression
An Information Theory of Compute-Optimal Size Scaling, Emergence, and Plateaus in Language Models
Benchmarking neural lossless compression algorithms on multi-purpose astronomical image data
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
Breaking Smoothness: The Struggles of Neural Compressors with Discontinuous Mappings
Bridging the Gap between Diffusion Models and Universal Quantization for Image Compression
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
Communication Compression for Tensor Parallel LLM Inference
Compressing Recurrent Neural Networks for FPGA-accelerated Implementation in Fluorescence Lifetime Imaging
Conditional Hallucinations for Image Compression
Copula-based Estimation of Continuous Sources for a Class of Constrained Rate-Distortion Functions
Deep Clustering with Associative Memories
Dense Backpropagation Improves Routing for Sparsely-Gated Mixture-of-Experts
Differentiable Attention
Diffusion Models With Learned Adaptive Noise
Distillation of Discrete Diffusion through Dimensional Correlations
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
EAMQ: Environment-based Adaptive Model Quantization on Federated Reinforcement Learning
Efficient and Robust Spike Ensemble Coding of Signals
Efficient Compression of Sparse Accelerator Data Using Implicit Neural Representations and Importance Sampling
Efficient Model Compression Techniques with FishLeg
Empirical Upper Bounds for Unstructured Sparsity in Compute-Efficient Language Modeling
EXAQ: Exponent Aware Quantization For LLMs Acceleration
Exploiting Temporal Priors for Efficient Real-time Compression and Feedback of Wireless Channels
FinerCut: Finer-grained Interpretable Layer Pruning for Large Language Models
Flexible image decoding in learned image compression
Formalizing Limits of Knowledge Distillation Using Partial Information Decomposition
Fused-Layer CNNs for Memory-Efficient Inference on Microcontrollers
FV-NeRV: Neural Compression for Free Viewpoint Videos
Getting free Bits Back from Rotational Symmetries in LLMs
Graph Transformation Augmentation for Contrastive Learning of Graph-Level Representation: An Initial Exploration
Grow to Compress? Efficient Training of Robust Networks on the Edge
How Many Does It Take to Prune a Network: Comparing One-Shot vs. Iterative Pruning Regimes
Improving Knowledge Distillation with Teacher's Explanation
Information-theoretic Generalization Analysis for Vector-Quantized VAEs
Integration of Large Vision Models in Driver Monitoring Systems: Compressing and Distilling for Real-Time Automotive Applications
Interactions Across Blocks in Post-Training Quantization of Large Language Models
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations
Large Language Model Compression with Neural Architecture Search
Latent Probabilistic Dataset Distillation with Theoretical Guarantees
Layer-Importance guided Adaptive Quantization for Efficient Speech Emotion Recognition
Layer-wise Quantization for Distributed Variational Inequalities
Learnable Fourier-based Activations for Implicit Signal Representations
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization
LLM Vocabulary Compression for Low-Compute Environments
LORC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
Losslessly Compressible Neural Network Parameters
LSH-E Tells You What To Discard: An Adaptive Locality-Sensitive Strategy for KV Cache Compression
M2M-TAG: Training-Free Many-to-Many Token Aggregation for Vision Transformer Acceleration
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
MAPLE: Memory-Aware Predict and Load for Efficient LLM Inference
MCUCoder: Adaptive Bitrate Learned Video Compression for IoT Devices
Mind the Gap Between Synthetic and Real: Probing Transfer Capabilities of Stable Diffusion Images
Neural Compression for Multispectral Satellite Images
Neural Normalized Compression Distance and the Disconnect Between Compression and Classification
Non-interactive Remote Coordination
On the Relationship Between Model Training Dynamics and Early Pruning Periods
P-SpikeSSM: Harnessing Probabilistic Spiking State Space Models for Long-Range Dependency Tasks
Partially Frozen Random Networks Contain Compact Strong Lottery Tickets
Perception Loss Function Adaptive to Rate for Learned Video Compression
PerCo (SD): Open Perceptual Compression
Polar Codes for Channel Simulation
Prechastic Coding: An Alternative Approach to Neural Network Description Lengths
QIANets: Quantum-Integrated Adaptive Networks for Reduced Latency and Improved Inference Times in CNN Models
Randomly Pivoted V-optimal Design: Fast Data Selection under Low Intrinsic Dimension
Sample Compression Hypernetworks: From Generalization Bounds to Meta-Learning
Sample compression unleashed : New generalization bounds for real valued losses
SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
Self-Data Distillation for Recovering Quality in Pruned Large Language Models
Shrinking the Size of Deep Extreme Multi-Label Classification
Simple LLM Compression Recovery Using Dynamic Prompting with Theoretical Analysis
SNeRV: Scalable Neural Representations for Video Coding
SpikingVTG: Saliency Feedback Gating Enabled Spiking Video Temporal Grounding
Sustainable AI: Efficient Pruning of Large Language Models in Resource-Limited Environments
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
The Rate-Distortion-Perception Trade-Off with Algorithmic Realism
The Trichromatic Strong Lottery Ticket Hypothesis: Neural Compression With Three Primary Supermasks
Towards Scalable Compression with Universally Quantized Diffusion Models
Training Block-wise Sparse Models Using Kronecker Product Decomposition
Training-Free Visual Token Compression via Delayed Spatial Merging
Transformers Learn to Compress Variable-order Markov Chains in-Context
Unified Lookup Tables: Privacy-Preserving Foundation Models
Unifying Subsampling Pattern Variations for Compressed Sensing MRI with Neural Operators
Vector Quantization with Sorting Transformation
VRVQ: Variable Bitrate Residual Vector Quantization for Audio Compression
Wasserstein Distortion with Intrinsic $\sigma$-Maps
Weight-Sharing Method for Upsampling Layer from Feature Embedding Recursive Block
What Makes for Good Image Captions?