Self-Supervised Skill Optimization
Quick Answer
This paper shows that Self-Supervised Skill Optimization (SSO) introduces a framework for optimizing skills in large language models (LLMs) without ground-truth feedback.
Quick Take
By leveraging unlabeled task instances, SSO outperforms existing prompt optimizers in both closed-ended and open-ended tasks, achieving results comparable to GT-based optimizers. This method enhances the reusability of procedural guidance for agents.
Key Points
- SSO learns reusable skills solely from unlabeled task instances.
- It generates skill probes using a subset of execution results.
- An LLM judge evaluates behaviors without ground-truth labels.
- SSO outperforms GT-free optimizers on closed and open-ended tasks.
- The method achieves results comparable to GT-based optimizers.
DeepSignal Analysis
What happened
Self-Supervised Skill Optimization (SSO) is a new framework designed to enhance skills in large language models (LLMs) without relying on ground-truth feedback. It utilizes unlabeled task instances to optimize skills, achieving performance levels comparable to ground-truth-based optimizers. SSO demonstrates effectiveness in both closed-ended and open-ended tasks.
Key evidence
- SSO operates by running current skills on unlabeled batches and generating skill probes from the results, which are then evaluated by an LLM judge.
- The framework aggregates evidence for and against observed behaviors across instances, ranking them to create a new skill that is accepted only if it outperforms the previous one on an unlabeled validation set.
- In tests, SSO outperformed existing ground-truth-free prompt optimizers and approached or exceeded the performance of the best ground-truth-based skill optimizer.
Why it matters
The introduction of SSO is significant as it addresses the challenge of optimizing LLM skills in scenarios where ground-truth feedback is unavailable. This advancement could enhance the utility of LLMs in various applications, making them more adaptable and efficient. By improving skill optimization without needing labeled data, SSO could lower the barrier to deploying LLMs in real-world tasks.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Siran Peng, Cuiyu Yang, Tianyu Fu, Tianshuo Zhang, Haoyuan Zhang, Weisong Zhao, Anyang Su, Minghui Wu, Huiying Li, Xiangyu Zhu, Chenxu Zhao, Zhen Lei
Abstract:Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feedback. Many applications, however, lack GT labels, task scores, rewards, or reliable task-specific evaluators. We therefore introduce Self-Supervised Skill Optimization (SSO), a comparative framework that learns a reusable skill from unlabeled task instances alone. At each step, SSO runs the current skill on an unlabeled batch, uses a subset of the resulting executions to generate complete skill probes, and runs the probes on the same batch. An LLM judge compares the resulting answers, trajectories, artifacts, or terminal states. A separate behavior extractor identifies behavioral differences without seeing the judge's decisions. SSO uses these decisions to aggregate evidence for and against the observed behaviors across instances. It then ranks the behaviors by the resulting evidence and renders a new complete skill from the highest-ranked behaviors. The update is accepted only if the new skill outperforms the current one on an unlabeled validation set. SSO outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks. On closed-ended benchmarks, it approaches and sometimes exceeds the strongest GT-based skill optimizer without using any GT feedback.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.28777 [cs.CL] |
| (or arXiv:2607.28777v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28777 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Siran Peng [view email]
[v1]
Thu, 30 Jul 2026 19:04:34 UTC (454 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.