Будь ласка, використовуйте цей ідентифікатор, щоб цитувати або посилатися на цей матеріал:
https://repository.hneu.edu.ua/handle/123456789/41444Повний запис метаданих
| Поле DC | Значення | Мова |
|---|---|---|
| dc.contributor.author | Знахур С. | - |
| dc.date.accessioned | 2026-07-20T09:17:41Z | - |
| dc.date.available | 2026-07-20T09:17:41Z | - |
| dc.date.issued | 2026 | - |
| dc.identifier.citation | Знахур С. Neural network optimisation: quantisation techniques, practical benchmarking, and deployment recommendations / С. Знахур // Proceedings of the 4th International Scientific and Practical Conference "Scientific Research: Emerging Theories and Practical Breakthroughs", July 6–8, 2026, Edinburgh, Scotland. – Edinburgh : European Open Science Space, 2026. – P. 64-75. | uk_UA |
| dc.identifier.uri | https://repository.hneu.edu.ua/handle/123456789/41444 | - |
| dc.description.abstract | The paper addresses the deployment of deep convolutional neural networks under the memory and energy limitations typical of mobile hardware, focusing on quantisation as a practical optimisation technique. The author performs a systematic empirical comparison of five convolutional architectures (MobileNetV2, ResNet-50, DenseNet-121, EfficientNet-B3, EfficientNet-B4), each converted with TensorFlow Lite Converter into three numerical-precision formats — FP32 baseline, FP16 and INT8 — and benchmarked on a physical iPhone 13 device (Apple A15 Bionic) through a purpose-built React Native 0.82 application using the react-native-fast-tflite library. Classification accuracy is assessed on a 500-image ImageNet-v2 subset. The results demonstrate that FP16 conversion halves model size while keeping Top-1 accuracy loss below 0.3% for every architecture, whereas INT8 reduces file size by 71–74% and improves inference latency by up to 2.8 times for mobile-friendly networks, but yields no consistent latency gain for EfficientNet-B3/B4 because part of their computation graph falls back to CPU execution within the TensorFlow Lite Metal delegate. Three factors governing quantisation effectiveness are identified: the structural compatibility of an architecture with the hardware delegate, the representativeness of the calibration dataset, and the delegate's operator coverage. Based on these findings, the author formulates differentiated, evidence-based deployment recommendations for mobile AI developers, organised by architecture family and use-case requirements, and proposes a reusable conversion-and-benchmarking methodology applicable beyond the five architectures studied. | uk_UA |
| dc.language.iso | en | uk_UA |
| dc.subject | neural network quantisation | uk_UA |
| dc.subject | mobile deep learning | uk_UA |
| dc.subject | TensorFlow Lite | uk_UA |
| dc.subject | FP16 | uk_UA |
| dc.subject | INT8 | uk_UA |
| dc.subject | MobileNetV2 | uk_UA |
| dc.subject | ResNet-50 | uk_UA |
| dc.subject | DenseNet-121 | uk_UA |
| dc.subject | EfficientNet | uk_UA |
| dc.subject | on-device inference | uk_UA |
| dc.subject | React Native | uk_UA |
| dc.subject | Apple A15 Bionic | uk_UA |
| dc.title | Neural network optimisation: quantisation techniques, practical benchmarking, and deployment recommendations | uk_UA |
| dc.type | Article | uk_UA |
| Розташовується у зібраннях: | Статті (ІС) | |
Файли цього матеріалу:
| Файл | Опис | Розмір | Формат | |
|---|---|---|---|---|
| Serhii_Znakhur_Edinburg.pdf | 2,62 MB | Adobe PDF | Переглянути/відкрити |
Усі матеріали в архіві електронних ресурсів захищені авторським правом, всі права збережені.