![]() |
| Hadoop Data Balancing |
Hadoop Balancer:
This is tool provided
to balance the disk uses throughout the Hadoop cluster. I may happen sometime
that some of the nodes in the cluster becomes over utilized or underutilized,
which occurs due to addition of new nodes where newly added nodes may be
underutilized or if there are less number of nodes result in overutilization.
We can run balancer from more than 1 machine in the cluster to increase the
speed of balancing but it will increase bandwidth uses to very high.
This tool requires administrator
right on the Hadoop cluster to run.
Syntax of the
balancer:
bin/start-balancer.sh
[-threshold <threshold>]
Where
start-balancer.sh files resides in the bin directory of the Hadoop folder. And the
threshold is the parameter which decides target of balance, this lies in
fraction between 0,1 the default value is 10% if nothing is passed as the threshold
value.
This process does the transferring
of blocks between the nodes resulting network activity and if a production
cluster must be used cautiously, as it result in some block missing error or late
reply from the cluster.
This process can be
stopped any time if required using following command:
