MORSE
Agentic Multi-objective Molecular Optimization via Dynamic Routing of Property-specific Editor Networks
Note: This article was drafted with AI assistance and reviewed by Jin Hyuk Kim.
Overview
- Task: Multi-objective molecular optimization (MMO) aims to improve several competing chemical properties simultaneously, such as blood-brain barrier permeability and binding activity, while preserving the original molecular structure.
- Limitation of prior work: Existing LLM-based approaches attempt to satisfy all requested property changes in a single generation step, placing a heavy burden on a single decoding pass.
- Our idea: MORSE reframes MMO as a sequential decision process: an LLM-based router repeatedly chooses which property-specific editor to apply next, or when to stop, based on the current state of the molecule.
- Venue: MORSE was presented as a poster at AI4Sci Korea 2026 (Sep 28 – Oct 1, 2026, Seoul Dragon City) by Jin Hyuk Kim and Jonghwan Choi of Hallym University.
1. Method
MORSE has two trainable parts: a set of property-specific editors that propose local edits, and an LLM-based router that decides which editor to apply at each step.
1.1 Property-specific Editors
- Each target property (BBBP, DRD2, plogP, and QED) has a dedicated structure-constrained molecular VAE editor.
- All editors share a single property-agnostic foundation model (a Transformer-based VAE), while a separate LoRA module [1] provides property-specific specialization for each editor.
Phase 1: Pretraining
- The property-agnostic foundation model is pretrained on the ChEMBL database, first using a molecular reconstruction objective and then using structure-constrained generation objectives.
- The resulting checkpoint is shared and kept frozen across all property-specific editors.
Phase 2: SFT w/ LoRA
- For each target property, a separate LoRA module is applied to the decoder of the frozen foundation model.
- Each LoRA module undergoes supervised fine-tuning (SFT) on property-specific molecular pairs.
Phase 3: RL w/ LoRA
- Each LoRA module is further optimized via KL-anchored reinforcement learning (RL).
- The property reward is balanced against the reference log-likelihood, so each editor improves its property without drifting from the SFT model.
1.2 Router
MDP Formulation
- State: The instruction is enriched with the current molecule, its Lipinski descriptors, and its current target-property scores, then encoded by a frozen Qwen2.5-7B-Instruct model [2]; the mean-pooled representation is the state \(s_t\).
- Action: One of the property-specific editors, or a termination action.
- Transition: The selected editor generates candidates from the current molecule, and the candidate with the highest reward (Section 1.3) becomes the next molecule \(m_{t+1}\).
- Termination: An episode ends when the router selects the termination action or reaches the maximum number of timesteps.
Training
- Only the MLP-based policy and value networks are trained; the LLM encoder and all editors stay frozen.
- The policy and value networks are optimized using proximal policy optimization (PPO) [3], without any ground-truth editor sequence.
- During both training and inference, the router samples an action from the categorical distribution predicted by the policy.
1.3 Reward
The molecular reward is a weighted sum of structural similarity to the initial molecule and the proportion of improved target properties:
\[R(m_t, m_0) = \begin{cases} 0, & \text{if } m_t = m_0 \text{ or } \mathrm{Sim}(m_t, m_0) < \delta, \\ \alpha \cdot \mathrm{Sim}(m_t, m_0) + (1 - \alpha) \cdot \mathrm{Imp}(m_t, m_0), & \text{otherwise.} \end{cases}\]- \(\mathrm{Sim}(m_t, m_0)\): Tanimoto similarity [4] between \(m_t\) and \(m_0\).
- \(\mathrm{Imp}(m_t, m_0)\): proportion of target properties improved in the requested directions relative to \(m_0\).
- \(\alpha\): balance between the two terms (0.5); \(\delta\): minimum similarity threshold (0.4).
- The difference in molecular reward between consecutive states is used as the step-wise reward for router training.
2. Main Results
Setup: MORSE is evaluated on three tasks from the MuMOInstruct benchmark [5], each combining three of four properties (BBBP, DRD2, plogP, QED), under both seen and unseen instruction templates.
| Task | Model | Seen | Unseen | ||||
|---|---|---|---|---|---|---|---|
| SR | Sim | SR×Sim | SR | Sim | SR×Sim | ||
| BDP | RePO | 0.206 | 0.569 | 0.117 | 0.198 | 0.572 | 0.113 |
| MORSE | 0.736 | 0.314 | 0.231 | 0.730 | 0.301 | 0.220 | |
| BDQ | RePO | 0.160 | 0.365 | 0.058 | 0.170 | 0.322 | 0.055 |
| MORSE | 0.718 | 0.239 | 0.172 | 0.748 | 0.230 | 0.172 | |
| BPQ | RePO | 0.274 | 0.509 | 0.140 | 0.242 | 0.596 | 0.144 |
| MORSE | 0.840 | 0.235 | 0.197 | 0.826 | 0.231 | 0.190 | |
Key Findings:
- MORSE consistently outperforms RePO in SR and SR×Sim across all three tasks under both seen and unseen instruction settings.
- The comparable performance under seen and unseen instruction templates suggests that the router generalizes beyond the phrasing encountered during training.
- Although RePO achieves higher Sim, its substantially lower SR indicates that structural preservation alone is insufficient for multi-objective optimization. MORSE instead accepts a moderate reduction in similarity to achieve a substantially higher success rate.
- The current results remain preliminary because the evaluation covers only one benchmark. Moreover, MORSE depends on the quality of its individual editors, and a weak editor can limit the overall performance of the framework.
References
[1] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv:2106.09685.
[2] Yang, A., Yu, B., Li, C., Liu, D., Huang, F., Huang, H., Jiang, J., Tu, J., Zhang, J., Zhou, J., Lin, J., Dang, K., Yang, K., Yu, L., Li, M., Sun, M., Zhu, Q., Men, R., He, T., Xu, W., Yin, W., Yu, W., Qiu, X., Ren, X., Yang, X., Li, X., Xu, Z., & Zhang, Z. (2025). Qwen2.5-1M technical report. arXiv:2501.15383.
[3] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv:1707.06347.
[4] Rogers, D., & Hahn, M. (2010). Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5), 742–754.
[5] Dey, V., Hu, X., & Ning, X. (2025). Gellm3o: Generalizing large language models for multi-property molecule optimization. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 25192–25221.
[6] Li, X., Zhou, Z., Li, Z., Yao, J., Rong, Y., Zhang, L., & Han, B. (2026). Reference-guided policy optimization for molecular optimization via LLM reasoning. arXiv:2603.05900.