Efficient small file management in Hadoop distributed file system for enhanced e-government services
Fredrick Ishengoma
What the paper says
Purpose This paper introduces the Efficient Small File Management Algorithm (ESFMA) to overcome the challenge of small file inefficiency of Hadoop distributed file system (HDFS) for e-government services. Design/methodology/approach ESFMA is designed with the following features: hierarchical metadata architecture, caching, block aggregation, prefetching and locality-aware data placement. These are intended to optimize NameNode memory usage, metadata handling, data block management, I/O and network performance. The algorithm was implemented in experiments on HDFS with real e-government small files. Findings The experiments showed that ESFMA saves 10% of NameNode memory, 12% of metadata requests, 3.8% of data block use, 15% of read latency, 17% of write latency and 10% of network traffic. Practical implications This study suggests that implementation of ESFMA has the potential to enable better e-government services in HDFS to be run efficiently and effectively. Originality/value This paper presents an algorithm for small file management in HDFS, filling an important need in improving service efficiency and performance in e-government services.
2 citations
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.25 × 0.4 = 0.10 |
| M · momentum | 0.55 × 0.15 = 0.08 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.