mirror of
https://codeberg.org/scip/kleingebaeck.git
synced 2025-12-17 04:21:00 +01:00
Compare commits
2 Commits
doc/add-di
...
fix/gocrit
| Author | SHA1 | Date | |
|---|---|---|---|
| 4602eccdf0 | |||
| 81de0d1a8f |
1
Makefile
1
Makefile
@@ -63,6 +63,7 @@ lint:
|
||||
|
||||
lint-full:
|
||||
golangci-lint run --enable-all --exclude-use-default --disable exhaustivestruct,exhaustruct,depguard,interfacer,deadcode,golint,structcheck,scopelint,varcheck,ifshort,maligned,nosnakecase,godot,funlen,gofumpt,cyclop,noctx,gochecknoglobals,paralleltest
|
||||
gocritic check -enableAll *.go
|
||||
|
||||
testfuzzy: clean
|
||||
go test -fuzz ./... $(ARGS)
|
||||
|
||||
43
README-de.md
43
README-de.md
@@ -222,49 +222,6 @@ Sowie alle Bilder.
|
||||
Das Format kann man mit der Variable `template` in der Konfiguration
|
||||
ändern. Die `example.conf` enthält ein Beispiel für das Standard Template.
|
||||
|
||||
## Verhalten des Tools
|
||||
|
||||
Es gibt einige Dinge über das Verhalten von kleingebäck, über die Du
|
||||
Bescheid wissen solltest:
|
||||
|
||||
- alle HTML Seiten und Bilder werden immer heruntergeladen
|
||||
- es wird ein (konfigurierbarer) Useragent verwendet
|
||||
- HTTP Cookies werden beachtet
|
||||
- bei Fehlern wird dreimal mit unterschiedlichem Abstand erneut
|
||||
versucht
|
||||
- Bilder Downloads laufen parallelisiert mit leicht unterschiedlichen
|
||||
zeitlichen Abständen ab
|
||||
- Gleich aussehende Bilder werden nicht überschrieben
|
||||
|
||||
Der letzte Punkt muss genauer erläutert werden:
|
||||
|
||||
Wenn man bei Kleinanzeigen.de eine Anzeige einstellt und Bilder
|
||||
postet, werden diese dort in ihrer Grösse reduziert (durch Kompression
|
||||
und Verkleinerung der Bilder usw.). Diese reduzierten Bilder werden
|
||||
dann von kleingebäck heruntergeladen. Falls Du Deine original Bilder
|
||||
behalten hast, kannst Du diese danach in das Backupverzeichnis
|
||||
kopieren. Bei einem erneuten kleingebäck-Lauf werden diese Bilder dann
|
||||
nicht überschrieben.
|
||||
|
||||
Wir verwenden dafür einen Algorythmus namens [distance
|
||||
hashing](https://github.com/corona10/goimagehash). Dieser Algorithmus
|
||||
prüft die Ähnlichkeit von Bildern. Diese können in ihrer Auflösung,
|
||||
Kompression, Farbtiefe und vielem mehr manipuliert worden sein und
|
||||
trotzdem als das "gleiche Bild" erkannt werden (wohlgemerkt nicht "das
|
||||
selbe": die Dateien sind durchaus unterschiedlich!). Bis zu einer
|
||||
Distance von 5 überschreiben wir keine Bilder, weil wir dann davon
|
||||
ausgehen, dass das lokal Vorhandene das Original ist.
|
||||
|
||||
Bitte beachte aber, dass dies KEIN Cachingmechanismus ist: die Bilder
|
||||
werden trotzdem immer alle heruntergeladen. Das muss so sein, da wir
|
||||
uns nicht die Dateinamen anschauen können, da kleinanzeigen.de diese
|
||||
nämlich zu Zahlen umbenennt. Und die Dateinamen können sich auch
|
||||
ändern, wenn der User in der Anzeige die Bilder umarrangiert hat.
|
||||
|
||||
Du kannst dieses Verhalten mit der Option **--force** ausschalten. Du
|
||||
kannst ausserdem mit der Option **--ignoreerrors** auch alle Fehler
|
||||
ignorieren, die beim Bilderdownload auftreten könnten.
|
||||
|
||||
## Documentation
|
||||
|
||||
Die Dokumentation kann man
|
||||
|
||||
42
README.md
42
README.md
@@ -207,48 +207,6 @@ variable. The supplied sample config contains the default template.
|
||||
|
||||
All images will be stored in the same directory.
|
||||
|
||||
## Tool Behavior
|
||||
|
||||
There are a bunch of things you might want to know about the behavior
|
||||
of the kleingebäck tool:
|
||||
|
||||
- all HTML pages and IMAGEs are always being downloaded
|
||||
- we use a (customizable) user agent
|
||||
- we respect HTTP cookies
|
||||
- in the case of an error, the tool does 3 retries, the time it waits
|
||||
between tries is longer for each retry
|
||||
- image download is parallized using small time differences to look
|
||||
more natural
|
||||
- same images are not being overwritten on subsequent download
|
||||
|
||||
|
||||
The latter needs to be elaborated a bit more:
|
||||
|
||||
If you publish an ad on kleinanzeigen.de and post images, those images
|
||||
will be reduced in size by the site (by compressing and down sizing
|
||||
them). This reduced images will be downloaded by kleingebäck. However,
|
||||
you may still own the original images and may want to put them into
|
||||
that backup directory so that you have all things for one ad together.
|
||||
|
||||
You can easily do that, because kleingebäck won't overwrite those
|
||||
original images. It uses something called a distance hash using
|
||||
[goimagehash](https://github.com/corona10/goimagehash). This
|
||||
algorithmus checks the similarity of images. If an image has been
|
||||
resized it is still very similar to the original one. We accept a
|
||||
maximum of a distance of 5, everything above leads to overwrite.
|
||||
|
||||
This works with resizes, cropped and otherwise manipulated images as
|
||||
long as the image still shows the original contents good enough.
|
||||
|
||||
Also note, that this is NOT a caching mechanism: the images will be
|
||||
downloaded anyway during each run. We also can't look at the file
|
||||
names because kleinanzeigen.de renames all images to numbers. And
|
||||
those might even change if the user re-arranges the images.
|
||||
|
||||
You can override this behavior using the **--force** option. Another
|
||||
option, **--ignoreerrors**, can be used to ignore all kinds of image
|
||||
errors.
|
||||
|
||||
## Documentation
|
||||
|
||||
You can read the documentation [online](https://github.com/TLINDEN/kleingebaeck/blob/main/kleingebaeck.pod) or locally once you have installed kleingebaeck with: `kleingebaeck --manual`.
|
||||
|
||||
17
SECURITY.md
Normal file
17
SECURITY.md
Normal file
@@ -0,0 +1,17 @@
|
||||
# Security Policy
|
||||
|
||||
## Supported Versions
|
||||
|
||||
Only the latest release is supported. If you find an issue (any
|
||||
issue!), please check with the latest release first.
|
||||
|
||||
## Reporting a Vulnerability
|
||||
|
||||
I don't agree with the "responsible disclosure" process most projects
|
||||
(and companies) work these days.
|
||||
|
||||
So, if you find a vulnerability of any kind, please just open an
|
||||
[issue](https://github.com/TLINDEN/kleingebaeck/issues). Please add
|
||||
all details required to reproduce the vulnerability. You won't be chased.
|
||||
|
||||
That's just all about it.
|
||||
2
ad.go
2
ad.go
@@ -73,7 +73,7 @@ func (ad *Ad) Incomplete() bool {
|
||||
}
|
||||
|
||||
func (ad *Ad) CalculateExpire() {
|
||||
if len(ad.Created) > 0 {
|
||||
if ad.Created != "" {
|
||||
ts, err := time.Parse("02.01.2006", ad.Created)
|
||||
if err == nil {
|
||||
ad.Expire = ts.AddDate(0, ExpireMonths, ExpireDays).Format("02.01.2006")
|
||||
|
||||
2
fetch.go
2
fetch.go
@@ -52,7 +52,7 @@ func NewFetcher(conf *Config) (*Fetcher, error) {
|
||||
}
|
||||
|
||||
func (f *Fetcher) Get(uri string) (io.ReadCloser, error) {
|
||||
req, err := http.NewRequest(http.MethodGet, uri, nil)
|
||||
req, err := http.NewRequest(http.MethodGet, uri, http.NoBody)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("failed to create a new HTTP request obj: %w", err)
|
||||
}
|
||||
|
||||
4
image.go
4
image.go
@@ -49,7 +49,7 @@ func (img *Image) LogValue() slog.Value {
|
||||
// holds all images of an ad
|
||||
type Cache []*goimagehash.ImageHash
|
||||
|
||||
func NewImage(buf *bytes.Reader, filename string, uri string) *Image {
|
||||
func NewImage(buf *bytes.Reader, filename, uri string) *Image {
|
||||
img := &Image{
|
||||
Filename: filename,
|
||||
URI: uri,
|
||||
@@ -134,7 +134,7 @@ func ReadImages(addir string, dont bool) (Cache, error) {
|
||||
reader := bytes.NewReader(data.Bytes())
|
||||
|
||||
img := NewImage(reader, filename, "")
|
||||
if err = img.CalcHash(); err != nil {
|
||||
if err := img.CalcHash(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
|
||||
26
main_test.go
26
main_test.go
@@ -334,14 +334,14 @@ type Adsource struct {
|
||||
}
|
||||
|
||||
// Render a HTML template for an adlisting or an ad
|
||||
func GetTemplate(adconfigs []AdConfig, adconfig AdConfig, htmltemplate string) string {
|
||||
func GetTemplate(adconfigs []AdConfig, adconfig *AdConfig, htmltemplate string) string {
|
||||
tmpl, err := tpl.New("template").Parse(htmltemplate)
|
||||
if err != nil {
|
||||
panic(err)
|
||||
}
|
||||
|
||||
var out bytes.Buffer
|
||||
if len(adconfig.ID) == 0 {
|
||||
if adconfig.ID == "" {
|
||||
err = tmpl.Execute(&out, adconfigs)
|
||||
} else {
|
||||
err = tmpl.Execute(&out, adconfig)
|
||||
@@ -376,15 +376,15 @@ func InitValidSources() []Adsource {
|
||||
ads := []Adsource{
|
||||
{
|
||||
uri: fmt.Sprintf("%s%s?userId=1", Baseuri, Listuri),
|
||||
content: GetTemplate(list1, empty, LISTTPL),
|
||||
content: GetTemplate(list1, &empty, LISTTPL),
|
||||
},
|
||||
{
|
||||
uri: fmt.Sprintf("%s%s?userId=1&pageNum=2", Baseuri, Listuri),
|
||||
content: GetTemplate(list2, empty, LISTTPL),
|
||||
content: GetTemplate(list2, &empty, LISTTPL),
|
||||
},
|
||||
{
|
||||
uri: fmt.Sprintf("%s%s?userId=1&pageNum=3", Baseuri, Listuri),
|
||||
content: GetTemplate(list3, empty, LISTTPL),
|
||||
content: GetTemplate(list3, &empty, LISTTPL),
|
||||
},
|
||||
}
|
||||
|
||||
@@ -392,7 +392,7 @@ func InitValidSources() []Adsource {
|
||||
for _, ad := range adsrc {
|
||||
ads = append(ads, Adsource{
|
||||
uri: fmt.Sprintf("%s/s-anzeige/%s/%s", Baseuri, ad.Slug, ad.ID),
|
||||
content: GetTemplate(nil, ad, ADTPL),
|
||||
content: GetTemplate(nil, &ad, ADTPL),
|
||||
})
|
||||
}
|
||||
|
||||
@@ -405,28 +405,28 @@ func InitInvalidSources() []Adsource {
|
||||
{
|
||||
// valid ad page but without content
|
||||
uri: fmt.Sprintf("%s/s-anzeige/empty/1", Baseuri),
|
||||
content: GetTemplate(nil, empty, EMPTYPAGE),
|
||||
content: GetTemplate(nil, &empty, EMPTYPAGE),
|
||||
},
|
||||
{
|
||||
// some random foreign webpage
|
||||
uri: INVALIDURI,
|
||||
content: GetTemplate(nil, empty, "<html>foo</html>"),
|
||||
content: GetTemplate(nil, &empty, "<html>foo</html>"),
|
||||
},
|
||||
{
|
||||
// some invalid page path
|
||||
uri: fmt.Sprintf("%s/anzeige/name/1", Baseuri),
|
||||
content: GetTemplate(nil, empty, "<html></html>"),
|
||||
content: GetTemplate(nil, &empty, "<html></html>"),
|
||||
},
|
||||
{
|
||||
// some none-ad page
|
||||
uri: fmt.Sprintf("%s/anzeige/name/1/foo/bar", Baseuri),
|
||||
content: GetTemplate(nil, empty, "<html>HTTP 404: /eine-anzeige/ does not exist!</html>"),
|
||||
content: GetTemplate(nil, &empty, "<html>HTTP 404: /eine-anzeige/ does not exist!</html>"),
|
||||
status: 404,
|
||||
},
|
||||
{
|
||||
// valid ad page but 503
|
||||
uri: fmt.Sprintf("%s/s-anzeige/503/1", Baseuri),
|
||||
content: GetTemplate(nil, empty, "<html>HTTP 503: service unavailable</html>"),
|
||||
content: GetTemplate(nil, &empty, "<html>HTTP 503: service unavailable</html>"),
|
||||
status: 503,
|
||||
},
|
||||
}
|
||||
@@ -465,7 +465,7 @@ func SetIntercept(ads []Adsource) {
|
||||
}
|
||||
}
|
||||
|
||||
func VerifyAd(advertisement AdConfig) error {
|
||||
func VerifyAd(advertisement *AdConfig) error {
|
||||
body := advertisement.Title + advertisement.Price + advertisement.ID + "Kleinanzeigen => " +
|
||||
advertisement.Category + advertisement.Condition + advertisement.Created
|
||||
|
||||
@@ -525,7 +525,7 @@ func TestMain(t *testing.T) {
|
||||
|
||||
// verify if downloaded ads match
|
||||
for _, ad := range adsrc {
|
||||
if err := VerifyAd(ad); err != nil {
|
||||
if err := VerifyAd(&ad); err != nil {
|
||||
t.Errorf(err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user