From Pixels to PCells: A Neurosymbolic Approach to Photonic Component Creation
Quick Answer
PixCell is a neurosymbolic system that transforms visual photonic components into executable parametric programs, achieving a mean IoU of over 0.9, significantly outperforming traditional models.
Quick Take
The system demonstrates effective training of the Qwen3.6-35B-A3B model, improving IoU from 0.422 to 0.491 after guided revisions, establishing a robust framework for photonic component design.
Key Points
- PixCell achieves a mean IoU of 0.9+, with scores up to 0.974 across eight targets.
- The system enables cheaper deterministic visual verification compared to generation attempts.
- Qwen3.6-35B-A3B model's IoU improved from 0.422 to 0.491 after three revision rounds.
- PixCell reconstructs primitive programs for various photonic stack configurations.
- The framework supports training without supervised demonstrations, enhancing design efficiency.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We present PixCell, a neurosymbolic system in which multimodal agents convert a visually presented photonic component into a parametric program over a small domain-specific language (DSL) of geometric primitives. A system enabling deterministic visual verification renders evaluation asymmetrically cheaper than the generation attempt. While models using multi-seed sampling and iterative revision reach a mean best-turn IoU of only 0.416, multimodal agents through PixCell's interface and verifier consistently exceed 0.9 mean IoU, with scores reaching 0.974 and 0.955 across eight component targets while also satisfying source contracts. These results demonstrate that frontier multimodal agents can reliably understand and render executable parametric representations from visual targets. Using these live parameters, cross-stack studies on an interferometer reconstruct primitive programs that satisfy an 8.0 nm free spectral range target and the original footprint constraint on modeled 220-nm SOI, 400-nm SiN, and 400-nm TFLN stacks. PixCell further carries a paper-derived splitter from visual reconstruction through SOI full-wave simulation, producing symmetric propagation and balanced outputs. Finally, the same executable verifier supplies a training reward and dataset used to train a Qwen3.6-35B-A3B model with LoRA and GRPO without supervised demonstrations. On eight training-excluded paper figures, its mean champion IoU rises from 0.422 after eight initial attempts to 0.491 after three verifier-guided revision rounds. These results therefore establish a controlled framework for measuring, retargeting, and improving visual-to-parametric photonic component design.
| Comments: | 16 pages, 13 figures, 3 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Optics (physics.optics) |
| Cite as: | arXiv:2608.00084 [cs.CV] |
| (or arXiv:2608.00084v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00084 arXiv-issued DOI via DataCite |
Submission history
From: Aadarsh Agarwal [view email]
[v1]
Wed, 29 Jul 2026 17:43:57 UTC (2,092 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.