<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>ICLR on Minghui Chen</title>
		<link>https://minghuichen.com/categories/iclr/</link>
		<description>Recent content in ICLR on Minghui Chen</description>
		<generator>Hugo</generator>
		<language>en</language>
		
		
		
		
			<lastBuildDate>Wed, 28 Jan 2026 00:00:00 +0000</lastBuildDate>
		
			<atom:link href="https://minghuichen.com/categories/iclr/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Textual Equilibrium Propagation for Deep Compound AI Systems</title>
				<link>https://minghuichen.com/publication/iclr_2026_tep/</link>
				<pubDate>Wed, 28 Jan 2026 00:00:00 +0000</pubDate>
				<guid>https://minghuichen.com/publication/iclr_2026_tep/</guid>
				<description>&lt;p&gt;&lt;strong&gt;Authors:&lt;/strong&gt; Minghui Chen, Wenlong Deng, James Zou, Han Yu, Xiaoxiao Li&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Published in:&lt;/strong&gt; Accepted to The Fourteenth International Conference on Learning Representations (&lt;strong&gt;ICLR 2026&lt;/strong&gt;)&lt;/p&gt;&#xA;&#xA;&#xA;&#xA;&#xA;&lt;h2 id=&#34;abstract&#34;&gt;Abstract&#xA;  &lt;a href=&#34;#abstract&#34;&gt;&lt;svg class=&#34;anchor-symbol&#34; aria-hidden=&#34;true&#34; height=&#34;26&#34; width=&#34;26&#34; viewBox=&#34;0 0 22 22&#34; xmlns=&#34;http://www.w3.org/2000/svg&#34;&gt;&#xA;      &lt;path d=&#34;M0 0h24v24H0z&#34; fill=&#34;currentColor&#34;&gt;&lt;/path&gt;&#xA;      &lt;path d=&#34;M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76.0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71.0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71.0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76.0 5-2.24 5-5s-2.24-5-5-5z&#34;&gt;&lt;/path&gt;&#xA;    &lt;/svg&gt;&lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Large language models (LLMs) are increasingly deployed as part of compound AI systems that coordinate multiple modules, such as retrievers, tools, and verifiers, over long-horizon workflows. Recent approaches that propagate textual feedback globally, such as TextGrad, make it feasible to optimize such pipelines, but we find that performance degrades as system depth grows.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Can Textual Gradient Work in Federated Learning?</title>
				<link>https://minghuichen.com/publication/iclr_2025_fedtextgrad/</link>
				<pubDate>Fri, 24 Jan 2025 00:00:00 +0000</pubDate>
				<guid>https://minghuichen.com/publication/iclr_2025_fedtextgrad/</guid>
				<description>&lt;p&gt;&lt;strong&gt;Authors:&lt;/strong&gt; Minghui Chen, Ruinan Jin, Wenlong Deng, Yuanyuan Chen, Zhi Huang, Han Yu, Xiaoxiao Li&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Published in:&lt;/strong&gt; The Thirteenth International Conference on Learning Representations (&lt;strong&gt;ICLR 2025&lt;/strong&gt;)&lt;/p&gt;&#xA;&#xA;&#xA;&#xA;&#xA;&lt;h2 id=&#34;abstract&#34;&gt;Abstract&#xA;  &lt;a href=&#34;#abstract&#34;&gt;&lt;svg class=&#34;anchor-symbol&#34; aria-hidden=&#34;true&#34; height=&#34;26&#34; width=&#34;26&#34; viewBox=&#34;0 0 22 22&#34; xmlns=&#34;http://www.w3.org/2000/svg&#34;&gt;&#xA;      &lt;path d=&#34;M0 0h24v24H0z&#34; fill=&#34;currentColor&#34;&gt;&lt;/path&gt;&#xA;      &lt;path d=&#34;M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76.0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71.0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71.0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76.0 5-2.24 5-5s-2.24-5-5-5z&#34;&gt;&lt;/path&gt;&#xA;    &lt;/svg&gt;&lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Recent studies highlight the promise of LLM-based prompt optimization, especially with TextGrad, which automates &amp;ldquo;differentiation&amp;rdquo; via texts and backpropagates textual feedback provided by LLMs. This approach facilitates training in various real-world applications that do not support numerical gradient propagation or loss calculation. It opens new avenues for optimization in decentralized, resource-constrained environments, suggesting that users of black-box LLMs (e.g., ChatGPT) could enhance components of LLM agentic systems (such as prompt optimization) through collaborative paradigms like federated learning (FL).&lt;/p&gt;</description>
			</item>
			<item>
				<title>Revisiting Delta-Parameter Pruning For Fine-Tuned Models</title>
				<link>https://minghuichen.com/publication/iclr_2025_xdare/</link>
				<pubDate>Thu, 23 Jan 2025 00:00:00 +0000</pubDate>
				<guid>https://minghuichen.com/publication/iclr_2025_xdare/</guid>
				<description>&lt;p&gt;&lt;strong&gt;Authors:&lt;/strong&gt; Wenlong Deng, Yize Zhao, Vala Vakilian, Minghui Chen, Xiaoxiao Li, Christos Thrampoulidis&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Published in:&lt;/strong&gt; The Thirteenth International Conference on Learning Representations (&lt;strong&gt;ICLR 2025, Spotlight&lt;/strong&gt;)&lt;/p&gt;&#xA;&#xA;&#xA;&#xA;&#xA;&lt;h2 id=&#34;abstract&#34;&gt;Abstract&#xA;  &lt;a href=&#34;#abstract&#34;&gt;&lt;svg class=&#34;anchor-symbol&#34; aria-hidden=&#34;true&#34; height=&#34;26&#34; width=&#34;26&#34; viewBox=&#34;0 0 22 22&#34; xmlns=&#34;http://www.w3.org/2000/svg&#34;&gt;&#xA;      &lt;path d=&#34;M0 0h24v24H0z&#34; fill=&#34;currentColor&#34;&gt;&lt;/path&gt;&#xA;      &lt;path d=&#34;M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76.0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71.0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71.0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76.0 5-2.24 5-5s-2.24-5-5-5z&#34;&gt;&lt;/path&gt;&#xA;    &lt;/svg&gt;&lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Storing open-source fine-tuned models separately introduces redundancy and increases response times in applications utilizing multiple models. Delta-parameter pruning (DPP), particularly the random drop and rescale (DARE) method proposed by Yu et al., addresses this by pruning the majority of delta parameters—the differences between fine-tuned and pre-trained model weights—while typically maintaining minimal performance loss. However, DARE fails when either the pruning rate or the magnitude of the delta parameters is large. We highlight two key reasons for this failure - (1) an excessively large rescaling factor as pruning rates increase, and (2) high mean and variance in the delta parameters. To address these, we develop two algorithmic improvements - (1) DARq, which modifies the rescaling factor in DARE, leading to significant performance gains at high pruning rates (e.g., &amp;gt;30% on COLA and SST2 for encoder models, with even larger improvements in decoder models), and (2) AdamR, an in-training modification that incorporates appropriate Delta regularization before applying DPP. We also demonstrate that DARq can be seamlessly combined with vanilla parameter-efficient fine-tuning techniques like LoRA and can facilitate structural DPP. Additionally, we revisit the application of importance-based pruning techniques within DPP, demonstrating that they outperform random-based methods when delta parameters are large. Through this comprehensive study, we develop a pipeline for selecting the most appropriate DPP method under various practical scenarios.&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
