NeurIPS 2025PastOther
NeurIPS 2025 Fourth Workshop on Deep Learning for Code
DL4C @ NeurIPS 2025
- Submission deadline
- Aug 28, 2025, 20:00 UTCimported from OpenReview — check the website for extensions
- Submission portal
- OpenReview
- Notes
- Topics were auto-suggested and may be imprecise — edits welcome.
Accepted papers (69)
Fetched from OpenReview (v2) on 2026-06-10.
A Matter of Representation: Towards Graph-Based Abstract Code Generation
A Note on the Code Quality Score System: LLMs for Maintainable Large Codebases
Adapting Language Models for Low-Resource Programming Languages
Advancing Environment Setup LLMs through Online Reinforcement Learning
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
Agint: Agentic Graph Compilation for Software Engineering Agents
Asm2SrcEval: Evaluating Large Language Models for Assembly to Source Code Translation
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
Beyond Accuracy: Realistic and Diagnostic Evaluation of Code Generation Models
BUILD-BENCH: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
Can Test-Time Compute Help LLMs Write Low-Resource Parallel Code Better?
ChopChop: Semantically Constraining the Code Output of Language Models
Code2Video: A Code-centric Paradigm for Educational Video Generation
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs
CoDyn: Dynamic LLM Routing for Coding Tasks
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
Cyber-Zero: Training Cybersecurity Agents without Runtime
Deep-Reproducer: From Paper Understanding to Code Generation
Demystify the Potential of Large Language Models as General-Purpose Surrogate Code Executors
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
Efficient Code Embeddings from Code Generation Models
Ensuring Functional Correctness of Large Code Models with Selective Generation
EquiBench: Benchmarking Large Language Models’ Understanding of Program Semantics via Equivalence Checking
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
Good-Enough Structured Generation: A Case Study on JSON Schema
HardTests: Synthesizing High-Quality Test Cases for LLM Coding
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
Improving Assembly Code Performance with Large Language Models via Reinforcement Learning
Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
In-Context Learning for Esoteric Programming Languages: Evaluating and Enhancing LLM Reasoning Without Fine-Tuning
Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models
Learning From Design Procedure To Generate CAD Programs for Data Augmentation
Learning to Solve and Verify: A Self-Play Framework for Mutually Improving Code and Test Generation
LLM-Driven Multi-step Translation from C to Rust using Static Analysis
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
Practical Code RAG at Scale: Task-Aware Retrieval Design Choices under Compute Budgets
pydra: Probing Code Representations With Synthetic Clones and Bugs
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
Random Baselines for Simple Code Problems are Competitive with Code Evolution
Refactoring Codebases through Library Design
RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
Security Knowledge Dilution in Large Language Models: How Irrelevant Context Degrades Critical Domain Expertise
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction
STACKFEED: Structured Textual Actor-Critic Knowledge base editing with FEEDback
SubtaskEval: Benchmarking LLMs on Competitive Programming Subtasks
SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
Thyme: Think Beyond Images
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
Training LLM Agents to Empower Humans
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
Where's the Bug? Attention Probing for Scalable Fault Localization
Workflows vs Agents for Code Translation