HIVE中UDTF編寫和使用

原創

spider_d

2020-02-24 19:09

1. UDTF介紹

UDTF(User-Defined Table-Generating Functions) 用來解決輸入一行輸出多行(On-to-many maping) 的需求。

2. 編寫自己需要的UDTF

繼承org.apache.hadoop.hive.ql.udf.generic.GenericUDTF。
實現initialize, process, close三個方法
UDTF首先會調用initialize方法，此方法返回UDTF的返回行的信息（返回個數，類型）。初始化完成後，會調用process方法，對傳入的參數進行處理，可以通過forword()方法把結果返回。最後close()方法調用，對需要清理的方法進行清理。

下面是我寫的一個用來切分”key:value;key:value;”這種字符串，返回結果爲key, value兩個字段。供參考：

import java.util.ArrayList;

import org.apache.hadoop.hive.ql.udf.generic.GenericUDTF;
import org.apache.hadoop.hive.ql.exec.UDFArgumentException;
import org.apache.hadoop.hive.ql.exec.UDFArgumentLengthException;
import org.apache.hadoop.hive.ql.metadata.HiveException;
import org.apache.hadoop.hive.serde2.objectinspector.ObjectInspector;
import org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorFactory;
import org.apache.hadoop.hive.serde2.objectinspector.StructObjectInspector;
import org.apache.hadoop.hive.serde2.objectinspector.primitive.PrimitiveObjectInspectorFactory;

public class ExplodeMap extends GenericUDTF{

    @Override
    public void close() throws HiveException {
        // TODO Auto-generated method stub    
    }

    @Override
    public StructObjectInspector initialize(ObjectInspector[] args)
            throws UDFArgumentException {
        if (args.length != 1) {
            throw new UDFArgumentLengthException("ExplodeMap takes only one argument");
        }
        if (args[0].getCategory() != ObjectInspector.Category.PRIMITIVE) {
            throw new UDFArgumentException("ExplodeMap takes string as a parameter");
        }

        ArrayList<String> fieldNames = new ArrayList<String>();
        ArrayList<ObjectInspector> fieldOIs = new ArrayList<ObjectInspector>();
        fieldNames.add("col1");
        fieldOIs.add(PrimitiveObjectInspectorFactory.javaStringObjectInspector);
        fieldNames.add("col2");
        fieldOIs.add(PrimitiveObjectInspectorFactory.javaStringObjectInspector);

        return ObjectInspectorFactory.getStandardStructObjectInspector(fieldNames,fieldOIs);
    }

    @Override
    public void process(Object[] args) throws HiveException {
        String input = args[0].toString();
        String[] test = input.split(";");
        for(int i=0; i<test.length; i++) {
            try {
                String[] result = test[i].split(":");
                forward(result);
            } catch (Exception e) {
                continue;
            }
        }
    }
}

3. 使用方法

UDTF有兩種使用方法，一種直接放到select後面，一種和lateral view一起使用。

1：直接select中使用：select explode_map(properties) as (col1,col2) from src;

不可以添加其他字段使用：select a, explode_map(properties) as (col1,col2) from src
不可以嵌套調用：select explode_map(explode_map(properties)) from src
不可以和group by/cluster by/distribute by/sort by一起使用：select explode_map(properties) as (col1,col2) from src group by col1, col2
2：和lateral view一起使用：select src.id, mytable.col1, mytable.col2 from src lateral view explode_map(properties) mytable as col1, col2;

此方法更爲方便日常使用。執行過程相當於單獨執行了兩次抽取，然後union到一個表裏。

4. 參考文檔

http://wiki.apache.org/hadoop/Hive/LanguageManual/UDF

http://wiki.apache.org/hadoop/Hive/DeveloperGuide/UDTF

http://www.slideshare.net/pauly1/userdefined-table-generating-functions

發表評論

所有評論

還沒有人評論，想成為第一個評論的人麼? 請在上方評論欄輸入並且點擊發布.

HIVE中UDTF編寫和使用

java中volatile關鍵字的含義

設計模式之---單例模式

sql和hql的區別

Java 中Thread用法

python中的itertools模塊

Mac下配置sublime實現LaTeX

https://yachay.unat.edu.pe/blog/index.php?comment_area=format_blog&comment_component=blog&comment_co

linux以太網驅動總結