MORSE

Agentic Multi-objective Molecular Optimization via Dynamic Routing of Property-specific Editor Networks

Note: This article was drafted with AI assistance and reviewed by Jin Hyuk Kim.

Overview

  • Task: Multi-objective molecular optimization (MMO) aims to improve several competing chemical properties simultaneously, such as blood-brain barrier permeability and binding activity, while preserving the original molecular structure.
  • Limitation of prior work: Existing LLM-based approaches attempt to satisfy all requested property changes in a single generation step, placing a heavy burden on a single decoding pass.
  • Our idea: MORSE reframes MMO as a sequential decision process: an LLM-based router repeatedly chooses which property-specific editor to apply next, or when to stop, based on the current state of the molecule.
  • Venue: MORSE was presented as a poster at AI4Sci Korea 2026 (Sep 28 – Oct 1, 2026, Seoul Dragon City) by Jin Hyuk Kim and Jonghwan Choi of Hallym University.

1. Method

Fig. 1. Overview of the proposed framework. At each timestep, an LLM-based router either selects a property-specific editor or terminates the optimization based on the current molecule. The selected editor produces an intermediate molecule, which is then used in the next routing decision.

MORSE has two trainable parts: a set of property-specific editors that propose local edits, and an LLM-based router that decides which editor to apply at each step.

1.1 Property-specific Editors

  • Each target property (BBBP, DRD2, plogP, and QED) has a dedicated structure-constrained molecular VAE editor.
  • All editors share a single property-agnostic foundation model (a Transformer-based VAE), while a separate LoRA module [1] provides property-specific specialization for each editor.

Phase 1: Pretraining

  • The property-agnostic foundation model is pretrained on the ChEMBL database, first using a molecular reconstruction objective and then using structure-constrained generation objectives.
  • The resulting checkpoint is shared and kept frozen across all property-specific editors.

Phase 2: SFT w/ LoRA

  • For each target property, a separate LoRA module is applied to the decoder of the frozen foundation model.
  • Each LoRA module undergoes supervised fine-tuning (SFT) on property-specific molecular pairs.

Phase 3: RL w/ LoRA

  • Each LoRA module is further optimized via KL-anchored reinforcement learning (RL).
  • The property reward is balanced against the reference log-likelihood, so each editor improves its property without drifting from the SFT model.

1.2 Router

MDP Formulation

  • State: The instruction is enriched with the current molecule, its Lipinski descriptors, and its current target-property scores, then encoded by a frozen Qwen2.5-7B-Instruct model [2]; the mean-pooled representation is the state \(s_t\).
  • Action: One of the property-specific editors, or a termination action.
  • Transition: The selected editor generates candidates from the current molecule, and the candidate with the highest reward (Section 1.3) becomes the next molecule \(m_{t+1}\).
  • Termination: An episode ends when the router selects the termination action or reaches the maximum number of timesteps.

Training

  • Only the MLP-based policy and value networks are trained; the LLM encoder and all editors stay frozen.
  • The policy and value networks are optimized using proximal policy optimization (PPO) [3], without any ground-truth editor sequence.
  • During both training and inference, the router samples an action from the categorical distribution predicted by the policy.

1.3 Reward

The molecular reward is a weighted sum of structural similarity to the initial molecule and the proportion of improved target properties:

\[R(m_t, m_0) = \begin{cases} 0, & \text{if } m_t = m_0 \text{ or } \mathrm{Sim}(m_t, m_0) < \delta, \\ \alpha \cdot \mathrm{Sim}(m_t, m_0) + (1 - \alpha) \cdot \mathrm{Imp}(m_t, m_0), & \text{otherwise.} \end{cases}\]
  • \(\mathrm{Sim}(m_t, m_0)\): Tanimoto similarity [4] between \(m_t\) and \(m_0\).
  • \(\mathrm{Imp}(m_t, m_0)\): proportion of target properties improved in the requested directions relative to \(m_0\).
  • \(\alpha\): balance between the two terms (0.5); \(\delta\): minimum similarity threshold (0.4).
  • The difference in molecular reward between consecutive states is used as the step-wise reward for router training.

2. Main Results

Setup: MORSE is evaluated on three tasks from the MuMOInstruct benchmark [5], each combining three of four properties (BBBP, DRD2, plogP, QED), under both seen and unseen instruction templates.

Task Model Seen Unseen
SR Sim SR×Sim SR Sim SR×Sim
BDP RePO 0.206 0.569 0.117 0.198 0.572 0.113
MORSE 0.736 0.314 0.231 0.730 0.301 0.220
BDQ RePO 0.160 0.365 0.058 0.170 0.322 0.055
MORSE 0.718 0.239 0.172 0.748 0.230 0.172
BPQ RePO 0.274 0.509 0.140 0.242 0.596 0.144
MORSE 0.840 0.235 0.197 0.826 0.231 0.190
Table 1. Performance comparison between MORSE and RePO [6] on the MuMOInstruct benchmark.

Key Findings:

  • MORSE consistently outperforms RePO in SR and SR×Sim across all three tasks under both seen and unseen instruction settings.
  • The comparable performance under seen and unseen instruction templates suggests that the router generalizes beyond the phrasing encountered during training.
  • Although RePO achieves higher Sim, its substantially lower SR indicates that structural preservation alone is insufficient for multi-objective optimization. MORSE instead accepts a moderate reduction in similarity to achieve a substantially higher success rate.
  • The current results remain preliminary because the evaluation covers only one benchmark. Moreover, MORSE depends on the quality of its individual editors, and a weak editor can limit the overall performance of the framework.

References

[1] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv:2106.09685.

[2] Yang, A., Yu, B., Li, C., Liu, D., Huang, F., Huang, H., Jiang, J., Tu, J., Zhang, J., Zhou, J., Lin, J., Dang, K., Yang, K., Yu, L., Li, M., Sun, M., Zhu, Q., Men, R., He, T., Xu, W., Yin, W., Yu, W., Qiu, X., Ren, X., Yang, X., Li, X., Xu, Z., & Zhang, Z. (2025). Qwen2.5-1M technical report. arXiv:2501.15383.

[3] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv:1707.06347.

[4] Rogers, D., & Hahn, M. (2010). Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5), 742–754.

[5] Dey, V., Hu, X., & Ning, X. (2025). Gellm3o: Generalizing large language models for multi-property molecule optimization. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 25192–25221.

[6] Li, X., Zhou, Z., Li, Z., Yao, J., Rong, Y., Zhang, L., & Han, B. (2026). Reference-guided policy optimization for molecular optimization via LLM reasoning. arXiv:2603.05900.