I have trouble with a web crawler using the TOR network. It's misusing the gopher proxy on my page. I don't want to disable/block tor (that would be the easy way out). It's permanently changing user agents and ignoring robots.txt. It ignores HTTP status codes. I'm currently serving it 4MB binary garbage in form of Link. It sucked in about 40GB of data now, but it doesn't explode and keeps crawling. Any other idea about what to do with it?
Phlog update: gopher://codevoid.de/0/posts/2019-04-27-manage-dotfiles-with-git.txt (https protocol works too)
@nblade: Check my tw.txt file. The specification does not allow a comment. I've added this now: 1970-01-01T01:00:00.000000Z▸FF:https://codevoid.de/tw.following.txt. I'd use the special date/time + FF: comment as trigger. This is backwards compatible and shouldn't really come up in anyones' timeline.
@nblade: It's just an idea. Not a clean one thoug, as clients would not know upfront who serves such a fiele and who not. Another idea would ne to mix a number of random followers into the twtxt file, which are updated when a person tweets.
The workflow app on iOS is magic. I now have a button that asks me to select a picture, then converts it to png, resizes it, strips the metadata, scps it to my jumphost, scps it further to my gopher jail and into my paste directory, constructs the http proxy URL and opens it in safari. All without user-interaction. Now I can share my mobile life with you guys! Prepare for cat pictures!
Why do most new cool things depend on hipster tech like nodejs? sigh #dat
Timeline Sandbox