An Intelligent Water Injection Decision-Making and Production Optimization Method Based on SAC Deep Reinforcement Learning
Received: 2026-04-21Accepted: 2026-05-22Published: 2026-06-25
AESIG 2026, 2(2), 17-33; https://doi.org/10.58244/aesig.263879
Abstract
To address the failure of static water injection strategies in late-stage oilfield development caused by strong reservoir heterogeneity, this paper proposes an intelligent water injection and dynamic production optimization method using Soft Actor-Critic (SAC) deep reinforcement learning. By formulating waterflooding optimization as a Markov Decision Process (MDP), a Finite Volume Method (FVM) simulator is dynamically coupled with a deep learning environment. Departing from traditional methods that rely solely on wellhead data, this model uses high-dimensional oil saturation images to capture the waterflood front’s topological evolution. A Convolutional Neural Network (CNN) extracts spatial features to output optimal real-time allocation weights for multiple injection wells, constrained by total injection volume. Tested on complex heterogeneous configurations—including “four-injector, five-producer” and “four-injector, nine-producer” patterns—the SAC agent demonstrated remarkable convergence stability and exploration efficiency. The model autonomously establishes an adaptive strategy that controls water cut, suppresses water channeling in high-permeability streaks, and intelligently redirects hydrodynamic energy to unswept zones. Compared to conventional uniform injection, this method significantly expands macroscopic sweep volume and reduces remaining oil saturation, offering a novel paradigm for the real-time, closed-loop management of complex reservoirs.
Keywords: Deep reinforcement learning; Soft Actor-Critic (SAC); Intelligent water injection decision-making; Closed-loop reservoir management; Spatial heterogeneity; Production optimization
Full Text
Download the full text PDF for viewing and using it according to the license of this paper.
Funding
This research was no funding provided.
References
- rouwer, D. R., & Jansen, J.-D. (2004). Dynamic Optimization of Waterflooding With Smart Wells Using Optimal Control Theory. SPE Journal, 9(04), 391–402. https://doi.org/10.2118/78278-PA
- Brunton, S. L., Noack, B. R., & Koumoutsakos, P. (2020). Machine Learning for Fluid Mechanics. Annual Review of Fluid Mechanics, 52(Volume 52, 2020), 477–508. https://doi.org/10.1146/annurev-fluid-010719-060214
- Cai, S., Mao, Z., Wang, Z., Yin, M., & Karniadakis, G. E. (2021). Physics-informed neural networks (PINNs) for fluid mechanics: A review. Acta Mechanica Sinica, 37(12), 1727–1738. https://doi.org/10.1007/s10409-021-01148-1
- Chen, Z., Zhang, K., Liu, P., Xin, G., Sun, Z., Tao, Z., Zhang, Y., Ji, W., Lu, Y., Jia, L., & Meng, H. (2025). Worst-Case Soft Actor-Critic-Based Safe Reinforcement Learning Method for Nonlinear Constrained Waterflood Reservoir Production Optimization. SPE Journal, 30(12), 7745–7766. https://doi.org/10.2118/230322-PA
- Ding, Y., Wang, X., Cao, X., Hu, H., & Bu, Y. (2023). A reinforcement learning method for optimal control of oil well production using cropped well group samples. Heliyon, 9(7). https://doi.org/10.1016/j.heliyon.2023.e17919
- Foroud, T., Baradaran, A., & Seifi, A. (2018). A comparative evaluation of global search algorithms in black box optimization of oil production: A case study on Brugge field. Journal of Petroleum Science and Engineering, 167, 131–151. https://doi.org/10.1016/j.petrol.2018.03.028
- Fujimoto, S., Hoof, H., & Meger, D. (2018). Addressing Function Approximation Error in Actor-Critic Methods. Proceedings of the 35th International Conference on Machine Learning, 1587–1596. https://proceedings.mlr.press/v80/fujimoto18a.html
- Garnier, P., Viquerat, J., Rabault, J., Larcher, A., Kuhnle, A., & Hachem, E. (2021). A review on deep reinforcement learning for fluid mechanics. Computers & Fluids, 225, 104973. https://doi.org/10.1016/j.compfluid.2021.104973
- Guo, Z., & Reynolds, A. C. (2018). Robust Life-Cycle Production Optimization With a Support-Vector-Regression Proxy. SPE Journal, 23(06), 2409–2427. https://doi.org/10.2118/191378-PA
- Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. Proceedings of the 35th International Conference on Machine Learning, 1861–1870. https://proceedings.mlr.press/v80/haarnoja18b.html
- Hourfar, F., Bidgoly, H. J., Moshiri, B., Salahshoor, K., & Elkamel, A. (2019). A reinforcement learning approach for waterflooding optimization in petroleum reservoirs. Engineering Applications of Artificial Intelligence, 77, 98–116. https://doi.org/10.1016/j.engappai.2018.09.019
- Hu, J., Ren, F., Wang, Z., & Jia, D. (2024). Efficient Scheduling of Constant Pressure Stratified Water Injection Flow Rate: A Deep Reinforcement Learning Method. IEEE Access, 12, 123856–123871. https://doi.org/10.1109/ACCESS.2024.3425837
- Jansen, J. D., Douma, S. D., Brouwer, D. R., Van den Hof, P. M. J., Bosgra, O. H., & Heemink, A. W. (2009, February 2). Closed-Loop Reservoir Management. SPE Reservoir Simulation Symposium. https://doi.org/10.2118/119098-MS
- Karniadakis, G. E., Kevrekidis, I. G., Lu, L., Perdikaris, P., Wang, S., & Yang, L. (2021). Physics-informed machine learning. Nature Reviews Physics, 3(6), 422–440. https://doi.org/10.1038/s42254-021-00314-5
- Lou, Q., Meng, X., & Karniadakis, G. E. (2021). Physics-informed neural networks for solving forward and inverse flow problems via the Boltzmann-BGK formulation. Journal of Computational Physics, 447, 110676. https://doi.org/10.1016/j.jcp.2021.110676
- Miftakhov, R., Al-Qasim, A., & Efremov, I. (2020, January 13). Deep Reinforcement Learning: Reservoir Optimization from Pixels. International Petroleum Technology Conference. https://doi.org/10.2523/IPTC-20151-MS
- Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. https://doi.org/10.1038/nature14236
- Mo, S., Zhu, Y., Zabaras, N., Shi, X., & Wu, J. (2019). Deep Convolutional Encoder-Decoder Networks for Uncertainty Quantification of Dynamic Multiphase Flow in Heterogeneous Media. Water Resources Research, 55(1), 703–728. https://doi.org/10.1029/2018WR023528
- Nasir, Y., & Durlofsky, L. J. (2023a). Deep reinforcement learning for optimal well control in subsurface systems with uncertain geology. Journal of Computational Physics, 477, 111945. https://doi.org/10.1016/j.jcp.2023.111945
- Nasir, Y., & Durlofsky, L. J. (2023b). Practical Closed-Loop Reservoir Management Using Deep Reinforcement Learning. SPE Journal, 28(03), 1135–1148. https://doi.org/10.2118/212237-PA
- Nwachukwu, A., Jeong, H., Pyrcz, M., & Lake, L. W. (2018). Fast evaluation of well placements in heterogeneous reservoir models using machine learning. Journal of Petroleum Science and Engineering, 163, 463–475. https://doi.org/10.1016/j.petrol.2018.01.019
- Rabault, J., Kuchta, M., Jensen, A., Réglade, U., & Cerardi, N. (2019). Artificial neural networks trained through deep reinforcement learning discover control strategies for active flow control. Journal of Fluid Mechanics, 865, 281–302. https://doi.org/10.1017/jfm.2019.62
- Rao, X. (2024a, November 4). The First Application of Quantum Computing Algorithm in Streamline-Based Simulation of Water-Flooding Reservoirs. ADIPEC. https://doi.org/10.2118/221850-MS
- Rao, X., Liu, Y., Fu, Q., He, X., Kwak, H., Zhao, H., & Hoteit, H. (2025, September 16). Boundary-Integral Type Neural Network (BINN) for Flow Problems in Anisotropic Reservoirs. Middle East Oil, Gas and Geosciences Show (MEOS GEO). https://doi.org/10.2118/227536-MS
- Rao, X., Liu, Y., He, X., & Hoteit, H. (2025). Physics-informed Kolmogorov–Arnold networks to model flow in heterogeneous porous media with a mixed pressure-velocity formulation. Physics of Fluids, 37(7), 076654. https://doi.org/10.1063/5.0279122
- Rao, X., Liu, Y., & Shen, Y. (2026). Quantum-Classical Physics-Informed Neural Networks for Solving Reservoir Seepage Equations (arXiv:2512.03923). arXiv. https://doi.org/10.48550/arXiv.2512.03923
- Rao, X., Luo, C., He, X., & Hyung, K. (2024, November 4). An Efficient Quantum Neural Network Model for Prediction of Carbon Dioxide CO2 Sequestration in Saline Aquifers. ADIPEC. https://doi.org/10.2118/222257-MS
- Rao, X. (饶翔). (2024b). Performance study of variational quantum linear solver with an improved ansatz for reservoir flow equations. Physics of Fluids, 36(4), 047104. https://doi.org/10.1063/5.0201739
- Ren, F., Rabault, J., & Tang, H. (2021). Applying deep reinforcement learning to active flow control in weakly turbulent conditions. Physics of Fluids, 33(3), 037121. https://doi.org/10.1063/5.0037371
- Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms (arXiv:1707.06347). arXiv. https://doi.org/10.48550/arXiv.1707.06347
- Tang, M., Liu, Y., & Durlofsky, L. J. (2020). A deep-learning-based surrogate model for data assimilation in dynamic subsurface flow problems. Journal of Computational Physics, 413, 109456. https://doi.org/10.1016/j.jcp.2020.109456
- Wang, Z.-Z., Zhang, K., Chen, G.-D., Zhang, J.-D., Wang, W.-D., Wang, H.-C., Zhang, L.-M., Yan, X., & Yao, J. (2023). Evolutionary-assisted reinforcement learning for reservoir real-time production optimization under uncertainty. Petroleum Science, 20(1), 261–276. https://doi.org/10.1016/j.petsci.2022.08.016
- Xin, G., Zhang, K., Wang, Z., Sun, Z., Zhang, L., Liu, P., Yang, Y., Sun, H., & Yao, J. (2024). Soft Actor-Critic Based Deep Reinforcement Learning Method for Production Optimization. In J. Lin (Ed.), Proceedings of the International Field Exploration and Development Conference 2023 (pp. 353–366). Springer Nature. https://doi.org/10.1007/978-981-97-0272-5_31
- Yan, B., & Zhong, Z. (2025). Deep reinforcement learning for optimal hydraulic fracturing design in real-time production optimization. Geoenergy Science and Engineering, 250, 213815. https://doi.org/10.1016/j.geoen.2025.213815
- Ye, M., & Elsheikh, A. H. (2025). Model-based reinforcement learning for active flow control. Physics of Fluids, 37(9), 093363. https://doi.org/10.1063/5.0287427
- Zhu, Y., & Zabaras, N. (2018). Bayesian deep convolutional encoder–decoder networks for surrogate modeling and uncertainty quantification. Journal of Computational Physics, 366, 415–447. https://doi.org/10.1016/j.jcp.2018.04.018


