Delete files older than 10days on HDFS

前端 未结 5 2077
难免孤独
难免孤独 2020-12-08 22:15

Is there a way to delete files older than 10 days on HDFS?

In Linux I would use:

find /path/to/directory/ -type f -mtime +10 -name \'*.txt\' -execdir         


        
5条回答
  •  -上瘾入骨i
    2020-12-08 22:42

    Solution 1: Using multiple commands as answered by daemon12

    hdfs dfs -ls /file/Path    |   tr -s " "    |    cut -d' ' -f6-8    |     grep "^[0-9]"    |    awk 'BEGIN{ MIN=14400; LAST=60*MIN; "date +%s" | getline NOW } { cmd="date -d'\''"$1" "$2"'\'' +%s"; cmd | getline WHEN; DIFF=NOW-WHEN; if(DIFF > LAST){ print "Deleting: "$3; system("hdfs dfs -rm -r "$3) }}'
    

    Solution 2: Using Shell script

    today=`date +'%s'`
    hdfs dfs -ls /file/Path/ | grep "^d" | while read line ; do
    dir_date=$(echo ${line} | awk '{print $6}')
    difference=$(( ( ${today} - $(date -d ${dir_date} +%s) ) / ( 24*60*60 ) ))
    filePath=$(echo ${line} | awk '{print $8}')
    
    if [ ${difference} -gt 10 ]; then
        hdfs dfs -rm -r $filePath
    fi
    done
    

提交回复
热议问题