win7 64位下安装hadoop的eclipse插件并编写运行WordCount程序

fgew2 · 发表于 2015-4-17 09:06:23

环境：
win7 64位
hadoop-2.6.0

步骤：

1、下载hadoop-eclipse-plugin-2.6.0.jar包
2、把hadoop-eclipse-plugin-2.6.0.jar放到eclipse安装目录下的plugins目录下
3、打开eclipse发现左边多出来一个DFS Locations

4、在win7上解压hadoop-2.6.0。
5、下载hadoop.dll、winutils.exe等文件。
要求支持hadoop2.6版本的（低版本的hadoop.dll会报错），让后拷贝下载文件到hadoop的bin目录，如果有已存在的文件直接跳过就行，不用覆盖原来的bin目录下的文件
6、设置Window->Prefrences->Hadoop Map/Reduce的installation directory为你解压的hadoop目录。

7、显示Map/Reduce选项卡。
选择Window->Open Perspective->Other->Map/Reduce

8、创建hdfs连接。
右键单击Map/Reduce Locations选项卡选择新建

9、弹出配置
这里面的Host、Port分别为你在mapred-site.xml、core-site.xml中配置的地址及端口

10、查看是否连接成功，是否可以看到hdfs上的文件

11、创建mapreduce程序

12、下一步填写项目名称OK。
13、如果自动导入了hadoop的jar包建议全部删除（因为hadoop-2.6.0不同目录下会有好多相同的jar包）都导进去就有重复的了。
14、搜索hadoop-2.6.0目录下所有的jar包拷贝到自己新建的一个目录，有重复的略过，然后把所有的jar包导入项目中。
15、编写wordcount程序新建一个class，内容如下：
package hdWordCount;
import java.io.IOException;
import java.util.StringTokenizer;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
import org.apache.hadoop.util.GenericOptionsParser;
public class WordCount {
public static class TokenizerMapper extends
Mapper {
  private final static IntWritable one = new IntWritable(1);
  private Text word = new Text();
  public void map(Object key, Text value, Context context)
throws IOException, InterruptedException {
StringTokenizer itr = new StringTokenizer(value.toString());
while (itr.hasMoreTokens()) {
word.set(itr.nextToken());
context.write(word, one);
}
  }
}
public static class IntSumReducer extends
Reducer {
  private IntWritable result = new IntWritable();
  public void reduce(Text key, Iterable values,
Context context) throws IOException, InterruptedException {
int sum = 0;
for (IntWritable val : values) {
sum += val.get();
}
result.set(sum);
context.write(key, result);
  }
}
public static void main(String[] args) throws Exception {
  Configuration conf = new Configuration();
  String[] otherArgs = new GenericOptionsParser(conf, args)
.getRemainingArgs();
  if (otherArgs.length != 2) {
System.err.println("Usage: WordCount  ");
System.exit(2);
  }
  Job job = new Job(conf, "word count");
  job.setJarByClass(WordCount.class);
  job.setMapperClass(TokenizerMapper.class);
  job.setCombinerClass(IntSumReducer.class);
  job.setReducerClass(IntSumReducer.class);
  job.setOutputKeyClass(Text.class);
  job.setOutputValueClass(IntWritable.class);
  FileInputFormat.addInputPath(job, new Path(otherArgs[0]));
  FileOutputFormat.setOutputPath(job, new Path(otherArgs[1]));
  System.exit(job.waitForCompletion(true) ? 0 : 1);
}
}

16、右键run as 选择配置 ,配置参数

一个为输入目录一个为输出目录（输出目录不能存在）

17、右键run as ，run on hadoop 完成，查看输出目录文件内容
18、完成！

账号		自动登录	找回密码
密码			立即注册

Centos6.5×64安装配置openmeetings3.0.3详

大疆运维招人啦，

C++ :try 语句块和异常处理

C++的多态

Red Hat RHCE 8 (EX294) Cert Guide

Java/C++ 区别：看完这一篇，就够用！

别再用过时库了！这 13 个顶级 C++ 库才是

[经验分享] win7 64位下安装hadoop的eclipse插件并编写运行WordCount程序

相关帖子

浏览过的版块

扫码加入运维网微信交流群