[转载]Linux下大文件的排序和去重复

原創

2019-04-07 03:21

Linux下大文件的排序和去重复

去重复行

　　简单的用法如下，如一个文件名：happybirthday.txt

　　cat happybirthday.txt (显示文件内容)

　　Happy Birthday to You!

　　Happy Birthday to You!

　　Happy Birthday Dear Tux!

　　Happy Birthday to You!

　　cat happybirthday.txt|sort （排序）

　　Happy Birthday Dear Tux!

　　Happy Birthday to You!

　　Happy Birthday to You!

　　Happy Birthday to You!

　　cat happybirthday.txt|sort|uniq (去重复行)

　　Happy Birthday Dear Tux!

　　Happy Birthday to You!

　　去大文件重复行

　　但有时碰到一个大文件时（例如G级的文件），用上面的命令时报错，提示空间不足。我尝试了一下，最后是用 split 命令把大文件分割为几个小文件，单独排完序后再合并 uniq 。

　　split -b 200m happybirthday.big Prefix_

　　用-b参数切割happybirthday.big，小文件为200M。切割后的文件名前缀是Prefix_切割后的文件名如

　　Prefix_aa

　　Prefix_ab再分别sort

　　sort Prefix_aa >Prefix_aa.sort

　　sort Prefix_ab >Prefix_ab.sort再用 sort -m合并，再 uniq

　　cat Prefix_aa.sort Prefix_ab.sort |sort -m |uniq这是好早前碰到的一个问题了。没记错的话应该是这么回事。~

　　sort 与 uniq 命令还有许多有用的参数，如sort -m、uniq -u、uniq -d等。sort 与 uniq的组合是很强大的。

發表評論

所有評論

還沒有人評論，想成為第一個評論的人麼? 請在上方評論欄輸入並且點擊發布.

相關文章

24-5-18 X

自 3 月 31 號回來之後，這兩個月像是失去了方向一般，對很多事都提不起興趣。今天和 X 聊了聊，他還是以前那個熟悉的樣子。高中的時候，他是我們班裏公認的第一，我們是普通班，但他有着實驗班的實力，事實上，高考時他全校第三，後來去了美國

Higurashi-kagome

2024-06-01 14:30:43

【dubbo】如何测试一个dubbo服务呢？

rpc服務框架——dubbo https://cn.dubbo.apache.org/zh-cn/blog/2023/02/23/一文幫你快速瞭解-dubbo-核心能力/ 自制項目： https://github.com/Jinwenxin

金大鑫要堅持

2024-06-01 14:29:53

kubeconfig 多个集群配置如何切换

kubectl config get-contexts kubectl config use-context <context-name> kubectl config current-context

2024-06-01 14:27:53

两台windowserver服务器配置Redis哨兵集群

十年河東，十年河西，莫欺少年窮學無止境，精益求精 redis下載地址：https://github.com/tporadowski/redis/releases 這裏選擇壓縮版，不選擇安裝版 1、集羣環境主機master: 局域網

2024-06-01 14:24:12

oidc-client.js踩坑吐槽贴

前言前面選用了IdentityServer4做爲認證授權的基礎框架,感興趣的可以看上篇<微服務下認證授權框架的探討>,已經初步完成了authorization-code與implicit的簡易demo(html+js 在IIS部署的站點)

2024-06-01 14:23:02

微盟电商-以造数工厂为底座的低成本自动化应用实现（一）

微盟電商-以造數工廠爲底座的低成本自動化應用實現 SAAS服務的特點是能夠以同一套代碼基礎，服務各種使用場景的客戶，由此帶來的業務組合與配置的多樣性是造成測試在造數環節以及自動化測試的實施階段面臨繁瑣與困難的根本原因。如何確保自動化的高效實

2024-06-01 14:20:12

Mac Brew install慢的问题

# 替換brew.git: jimmy@MacBook-Pro Library % cd "$(brew --repo)" jimmy@MacBook-Pro Homebrew % git remote set-url origin htt

2024-06-01 14:18:02

Vue devDependencies 与 dependencies 能别

Vue devDependencies 與 dependencies 能別，如何往項目的node_modules安裝組件概述 devDependencies 用於本地環境開發只會在開發環境下依賴的模塊，生產環境不會被打入包內（通過

2024-06-01 14:18:02

mysql 超大大数据库复制前可执行的加速导入的SQL

use 數據庫;set global innodb_flush_log_at_trx_commit=0;set global max_allowed_packet=1024*1024*20;set global bulk_insert_bu

2024-06-01 14:14:21

css25 CSS Tables

https://www.w3schools.com/css/css_table.asp css25 CSS Tables CSS Tables The look of an HTML table can be greatly improv

2024-06-01 14:13:21

css29 CSS Layout - The z-index Property

https://www.w3schools.com/css/css_z-index.asp CSS Layout - The z-index Property The z-index property specifies th

2024-06-01 14:13:21

css28 CSS Layout - The position Property

https://www.w3schools.com/css/css_positioning.asp CSS Layout - The position Property The position property specifies t

2024-06-01 14:13:21

css26 CSS Layout - The display Property

https://www.w3schools.com/css/css_display_visibility.asp CSS Layout - The display Property The display property is

2024-06-01 14:13:21

css31 CSS Layout - float and clear

https://www.w3schools.com/css/css_float.asp CSS Layout - float and clear The CSS float property specifies how an

2024-06-01 14:13:21

css27 CSS Layout - width and max-width

https://www.w3schools.com/css/css_max-width.asp CSS Layout - width and max-width Using width, max-width and margi

2024-06-01 14:13:21

24小時熱門文章

最新文章

最新評論文章