Research paper

A Real-Time Stitching Method for Dual-View Street-View Images Based on a TensorRT Hybrid Inference Architecture

By: Wengui Tian, Yawen Wang, Haoyu Zhao, Jiayu Bai
School of Computer Science and Engineering Xi’an Technological University Xi’an, China
Received: 2026-07-02Revised: 2026-08-04Accepted: 2026-09-10Published: 2026-09-22
IJANMC 2026, 11(4), 92-105; https://doi.org/10.58244/ijanmc.260007
This article is supported by “Xi'an Technology University's National-level College Students' Innovation and Entrepreneurship Training Program in 2025 (Project Number: 202510702022)”.

A Real-Time Stitching Method for Dual-View Street-View Images Based on a TensorRT Hybrid Inference Architecture

Abstract—Deep image stitching networks deliver strong alignment quality, yet their inference latency keeps them below the frame rate that dual-view street-view applications demand. This paper reports a segmented hybrid inference architecture built on TensorRT. The partition criterion is the numerical sensitivity of each sub-graph: compute-bound but numerically stable modules are handed to TensorRT, while sensitive modules stay in PyTorch. Using the UDIS2 unsupervised deep image stitching algorithm as its base, the architecture re-cuts the two-stage pipeline of the Warp geometric alignment network and the Composition fusion network into a compute-bound segment and a sensitivity-critical segment that run on different backends. The ResNet-50 backbone and the correlation computation module are exported as a TensorRT FP16 engine, which lets the half-precision units of the GPU Tensor Cores carry the bulk of the arithmetic. Direct linear transformation solving, homography warping and thin-plate spline interpolation remain in the PyTorch FP32 environment, so that matrix inversion never runs under a low-precision floating-point representation. The Composition fusion network is deployed as a separate TensorRT FP16 engine. On an NVIDIA RTX 4060 Laptop GPU the end-to-end latency falls from 213.9 ms for the original PyTorch pipeline to 39.7 ms, giving a frame rate of 25.2 FPS and a speedup of about 5.4. Fidelity at the Composition output reaches 51.3 dB peak signal-to-noise ratio, and a pixel-level audit of the final panorama places 99.88 percent of the canvas within one grey level of the gold-standard render. Speed and image quality are therefore obtained together rather than traded against each other.

Keywords-Deep Learning; Image Stitching; Tensorrt Acceleration; Mixed-Precision Inference; UDIS2; Real-Time Vision System

CC BY 4.0
© 2026 by author(s). Licensee MOSP, Macao, China. This is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY 4.0) license.
Disclaimer: The statements, opinions and data contained in this journal are solely those of the individual authors and contributors and not of Macao Scientific Publishers and/or the editors. Macao Scientific Publishers and/or the editors disclaim responsibility for any injury to persons or property resulting from any ideas, methods, instructions or products referred to in the content.
Cite

Reference format:

Wengui Tian, Yawen Wang, Haoyu Zhao, 等. A Real-Time Stitching Method for Dual-View Street-View Images Based on a TensorRT Hybrid Inference Architecture[J]. IJANMC, 2026, 11(4): 92-105.
Share

Copy the link below to share this article:

Contact Us

Contact us via:

Email
Not available
Telephone
Not available
Supplementary

No supplementary material is available for this article.

Download PDF
PDF1.7 MB

Enter code to download

Captcha

Submit Your Manuscript Now