NeurIPS 2024PastGenerative modelsFairness & ethicsMultimodal
Workshop on Responsibly Building the Next Generation of Multimodal Foundational Models
NeurIPS 2024 Workshop RBFM
- Submission deadline
- Sep 21, 2024, 23:59 UTCimported from OpenReview — check the website for extensions
- Submission portal
- OpenReview
- Notes
- Topics were auto-suggested and may be imprecise — edits welcome.
Accepted papers (34)
Fetched from OpenReview (v2) on 2026-06-10.
Adversarial Robust Deep Reinforcement Learning is Neither Robust Nor Safe
Aligning to What? Limits to RLHF Based Alignment
Attention Shift: Steering AI Away from Unsafe Content
BigDocs: An Open and Permissively-Licensed Dataset for Training Multimodal Models on Document and Code Tasks
Building and better understanding vision-language models: insights and future directions
Comparison Visual Instruction Tuning
Consistency-diversity-realism Pareto fronts of conditional image generative models
Coordinated Robustness Evaluation Framework for Vision Language Models
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
Decompose, Recompose, and Conquer: Multi-modal LLMs are Vulnerable to Compositional Adversarial Attacks in Multi-Image Queries
Exploring Intrinsic Fairness in Stable Diffusion
GUIDE: A Responsible Multimodal Approach for Enhanced Glaucoma Risk Modeling and Patient Trajectory Analysis
How to Determine the Preferred Image Distribution of a Black-Box Vision-Language Model?
Incorporating Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
Just rephrase it! Uncertainty estimation in closed-source language models via multiple rephrased queries
LEMoN: Label Error Detection using Multimodal Neighbors
LLAVAGUARD: VLM-based Safeguards for Vision Dataset Curation and Safety Assessment
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
Multimodal Situational Safety
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
Position Paper: Protocol Learning, Decentralized Frontier Risk and the No-Off Problem
Probabilistic Active Few-Shot Learning in Vision-Language Models
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
Skipping Computations in Multimodal LLMs
The Multi-faceted Monosemanticity in Multimodal Representations
Towards Secure and Private AI: A Framework for Decentralized Inference
Trust but Verify: Reliable VLM evaluation in-the-wild with program synthesis
When Do Universal Image Jailbreaks Transfer Between Vision-Language Models?
WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models
You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models