Будь ласка, використовуйте цей ідентифікатор, щоб цитувати або посилатися на цей матеріал:
https://repository.hneu.edu.ua/handle/123456789/41441| Назва: | Mobile information system based on optimized deep neural networks for resource-constrained platforms |
| Автори: | Знахур Л. |
| Теми: | mobile devices neural networks quantization fp16 int8 optimization pruning knowledge distillation tensorflow lite mobilenetv2 resnet-50 densenet-121 efficientnet-b3 efficientnet-b4 inference. |
| Дата публікації: | 2026 |
| Бібліографічний опис: | Знахур Л. Mobile information system based on optimized deep neural networks for resource-constrained platforms / Л. Знахур // Proceedings of the 4th International Scientific and Practical Conference "Scientific Research: Emerging Theories and Practical Breakthroughs", July 6–8, 2026, Edinburgh, Scotland. – Edinburgh : European Open Science Space, 2026. – P. 75-86. |
| Короткий огляд (реферат): | The paper is devoted to designing a mobile information system built around optimised deep neural networks intended for resource-constrained platforms. The author provides a comprehensive analysis of model-optimisation techniques — numerical quantisation, structural pruning and knowledge distillation — with a primary focus on post-training quantisation (PTQ) and quantisation-aware training (QAT), and formulates the underlying scale-factor and zero-point mapping equations used for INT8 conversion. The practical part experimentally evaluates five convolutional architectures — MobileNetV2, ResNet-50, DenseNet-121, EfficientNet-B3, EfficientNet-B4 — converted with the TensorFlow Lite framework and benchmarked with a single-threaded React Native CLI application on a physical Apple iPhone (A14 Bionic). Four metrics are analysed for every model and precision format: disk footprint, inference latency, classification accuracy, and resource efficiency. The results confirm that FP16 conversion halves storage requirements with a negligible accuracy loss of up to 0.1%, while full INT8 static quantisation reduces file size by up to 75% and accelerates inference by up to 62.1%, at the cost of accuracy degradation ranging from 0.6% for MobileNetV2 to 2.4% for EfficientNet-B4. This variation is linked to the structural properties of each network: depthwise separable convolutions and linear bottlenecks in MobileNetV2, together with the residual shortcuts of ResNet-50, provide strong resilience to fixed-point rounding errors, whereas the dense feature concatenation of DenseNet-121 and the tightly balanced compound scaling of the EfficientNet family make them considerably more sensitive to INT8 conversion. Part of the observed latency behaviour is attributed to the fallback of unsupported operators from INT8 kernels to FP32 CPU execution inside the TensorFlow Lite interpreter. Qualitative thermal observations, supported by published hardware benchmarks, additionally indicate 35–50% lower energy consumption for INT8 models compared with FP32 baselines. The findings support recommending quantisation-aware training for architecture families most sensitive to static INT8 quantisation, so as to preserve classification accuracy while retaining the compression and efficiency benefits required for mobile deployment. |
| URI (Уніфікований ідентифікатор ресурсу): | https://repository.hneu.edu.ua/handle/123456789/41441 |
| Розташовується у зібраннях: | Статті (ІС) |
Файли цього матеріалу:
| Файл | Опис | Розмір | Формат | |
|---|---|---|---|---|
| Liudmyla_Znakhur_Edinburg.pdf | 2,74 MB | Adobe PDF | Переглянути/відкрити |
Усі матеріали в архіві електронних ресурсів захищені авторським правом, всі права збережені.