Showing posts with label #Hadoop. Show all posts
Showing posts with label #Hadoop. Show all posts

Thursday, April 13, 2023

What happens when #hadoop is in #safemode

 When Hadoop is in safe mode, it means that the Hadoop NameNode has started up and is running, but is not yet ready to serve client requests. Safe mode is a protective mode that the NameNode enters automatically when it detects certain conditions in the cluster that require administrator attention. In safe mode, the following things happen:

  1. The NameNode does not allow any new file systems modifications, such as creating, deleting, or renaming files or directories.

  2. The NameNode periodically checks the status of each DataNode to make sure that it has the minimum number of replicas for each block. If a DataNode is missing a replica, the NameNode will begin replicating the missing block to another DataNode.

  3. The NameNode waits for a configurable threshold of DataNodes to report that they are alive and functioning properly. This threshold is configured using the dfs.safemode.min.datanodes property.

  4. The NameNode also waits for a configurable threshold of blocks to be reported as available by DataNodes. This threshold is configured using the dfs.safemode.threshold.pct property.

Once the NameNode has verified that the cluster is in a healthy state, it will exit safe mode and begin serving client requests. If any problems are detected during the safe mode period, the NameNode will remain in safe mode until the problems are resolved.

how #communication between #datanodes happens in #hadoop

 In Hadoop, the communication between DataNodes happens in the following way:

  1. Heartbeats: DataNodes periodically send heartbeats to the NameNode to indicate that they are alive and functioning properly. The frequency of these heartbeats can be configured by the administrator.

  2. Block Reports: DataNodes send block reports to the NameNode at startup and periodically afterward to report the list of blocks that they are currently storing. These block reports help the NameNode to maintain an up-to-date map of the cluster's data.

  3. Replication: When a DataNode detects that a block has become under-replicated (i.e., there are not enough copies of the block), it will request additional copies from other DataNodes that have replicas of the same block. This process is known as replication.

  4. Data Transfer: When a client wants to read or write a file, it first contacts the NameNode to get the location of the file's blocks. The client can then directly contact the DataNodes that are storing the blocks to read or write the data. The DataNodes communicate with each other to transfer data as needed to maintain the desired level of replication.

Tuesday, March 7, 2023

how to impliment #acl (#ACL : Access control list)in #hadoop

ACL (Access Control Lists) can be implemented in Hadoop to control access to HDFS (Hadoop Distributed File System) files and directories. Here are the steps to implement ACL in Hadoop:

  1. Enable ACL in HDFS by adding the following property to the hdfs-site.xml configuration file:
<property>
  <name>dfs.namenode.acls.enabled</name>
  <value>true</value>
</property>

  1. Restart the HDFS service to apply the configuration changes.

  2. Set ACL for a file or directory using the hdfs dfs -setfacl command. For example, to set ACL for a directory named "directory" and give the user "shashwat" read and write permissions, use the following command:

hdfs dfs -setfacl -m username:shashwat:rwx directory

  1. Check the ACL of a file or directory using the hdfs dfs -getfacl command. For example, to check the ACL of the "directory" directory, use the following command:
hdfs dfs -getfacl directory

This will display the ACL entries for the directory, including the permissions and users/groups assigned.

Note that ACL can also be set for groups and masks, in addition to users. The mask is used to restrict the maximum permissions that can be assigned to a file or directory. For example, if the mask is set to "r-x", then the maximum permission that can be assigned to a user/group is read and executed.

Also, keep in mind that ACLs only apply to HDFS, and do not affect access to other components in the Hadoop ecosystem such as MapReduce or YARN.



Thursday, March 2, 2023

how to add #proxy user in #oozie or #hue



 To add a proxy user in Oozie or Hue, you will need to update the configuration files for these tools.

Here are the steps to add a proxy user in Oozie:

1.       Open the core-site.xml file, which is located in the Hadoop configuration directory. This directory is typically located at /etc/hadoop/conf/.

2.       Add the following configuration property to the file, replacing proxyuser with the username of the user you want to act as a proxy:

<property>

  <name>hadoop.proxyuser.proxyuser.hosts</name>

  <value>*</value>

</property>

 

3.       Add the following configuration property to the file, replacing proxyuser with the username of the user you want to act as a proxy, and user with the username of the user you want to allow to use the proxy:

 

<property>

  <name>hadoop.proxyuser.proxyuser.groups</name>

  <value>user</value>

</property>

 

4.       Save and close the configuration file.

5.       Restart the Oozie server to apply the changes.

 

Here are the steps to add a proxy user in Hue:

 

1.       Open the hue.ini configuration file, located in the Hue configuration directory. This directory is typically located at /etc/hue/conf/.

2.       Find the [desktop] section in the configuration file, and add the following setting:

 

[desktop].

proxy_username=proxyuser

 

3.       Save and close the configuration file.

4.       Restart the Hue server to apply the changes.

 

Note that adding a proxy user allows a specific user to act on behalf of another user. This can be a security risk, so it is important to carefully consider which users are allowed to act as proxies and which users are allowed to use proxies.

Tuesday, February 28, 2023

Most common #apache #hadoop #error #messages

 1.       java.io.IOException: This error occurs when Hadoop encounters an issue while reading or writing data.

2.       File not found exception: This error occurs when Hadoop is unable to find the specified file or directory.

3.       NameNode is in Safe Mode: This error message indicates that the Hadoop NameNode is in safe mode, which restricts write operations to the Hadoop file system.

4.       Unable to create directory: This error occurs when Hadoop is unable to create a directory in the file system.

5.       Block Missing Exception: This error message indicates that a block of data is missing from the Hadoop file system.

6.       Permission denied: This error occurs when the user does not have the required permissions to perform the requested operation.

7.       Task attempt failed to report status: This error message indicates that the Hadoop job failed to report its status to the JobTracker.

8.       Exceeded maximum allowed attempts: This error occurs when a task in Hadoop exceeds the maximum number of allowed attempts.

9.       Namenode not starting: This error occurs when the Hadoop NameNode process fails to start, often due to an issue with the file system or configuration.

10.   DataNode not starting: This error occurs when the Hadoop DataNode process fails to start, often due to an issue with the file system or configuration.

11.   Corrupt block pool: This error occurs when the Hadoop NameNode detects corruption in the block pool, often due to hardware or file system issues.

12.   Incorrect block size: This error occurs when the Hadoop NameNode detects that a block has been written with an incorrect size, often due to a configuration issue or bug in the code.

13.   Permission denied: This error occurs when the user does not have the required permissions to perform the requested operation on the Hadoop file system.

14.   Invalid input: This error occurs when the input data provided to a Hadoop job is not valid or does not match the expected format.

15.   Connection refused: This error occurs when Hadoop is unable to connect to a remote service, often due to network issues or configuration problems.

16.   TaskTracker failed to start: This error occurs when the Hadoop TaskTracker process fails to start, often due to an issue with the configuration or file system.

Featured Posts

Explained: Complete Guide to systemd-logind, Power Keys, Laptop Lids, VTs & User Sessions

 If you use Linux on a laptop, desktop, server, or remote-access machine, there is a good chance that systemd-logind is quietly controlling...